AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Oct 1, 2026
·
policygovernanceai-safetyprocurementopenaianthropicgoogleevaluationsantitrustenterprise

The labs' private regulator has a name now — and the job it wants is deciding who counts as an independent auditor

TL;DR: On 24 September 2026 The Information reported that Google, OpenAI and Anthropic have agreed to form the Standards Authority for Frontier AI (SAFA) — a FINRA-style, industry-funded self-regulator operating outside government, targeting launch between end-2026 and early 2027. Reported remit: common standards and benchmark testing, support for third-party pre-deployment testing, incident reporting rules, and qualification requirements for independent auditors. Reported CEO shortlist: Sriram Krishnan, Arati Prabhakar, Condoleezza Rice, David Friedberg. Reported scientific consultants: Beth Barnes (METR) and Paul Christiano, who joined OpenAI’s Foundation Board on 8 September. Meta, xAI, Microsoft, Mistral, Cohere and every open-weight lab are not in it. The antitrust waiver Amodei asked for was blocked in the NDAA on 15 September; the labs are proceeding on a DOJ enforcement posture stated on 17 September instead. And all of this landed in the same twenty-four hours the White House asked OpenAI and Anthropic to withhold models from the UK AI Security Institute. The buyer consequence: “independently evaluated” is losing its referent. Ask for the artefacts, not the adjective.

What was reported

On 24 September 2026, The Information reported that Google, OpenAI and Anthropic have moved past working groups and agreed to establish an institution. Its working name is the Standards Authority for Frontier AI — SAFA. The target launch window is the end of this year or early 2027.

The structural model is FINRA: the Financial Industry Regulatory Authority, the industry-funded body that writes and enforces conduct rules for US brokerages under SEC registration. Applied to AI, that means a self-regulatory organisation funded by the firms it regulates, operating independently of government, requiring neither Congressional nor presidential approval to exist — though it would likely need to register with a federal agency to carry any enforceable authority.

The reported functions are four:

FunctionWhat it means
Common standards and benchmark testingSafety claims from different labs become comparable rather than self-graded
Support for third-party pre-deployment testingOutside groups get model access before release
Incident reporting rulesA defined channel and threshold for disclosing deployed-model failures
Auditor qualification requirementsSAFA decides who is permitted to count as an independent auditor

The first three are things enterprise buyers have asked for since 2024. The fourth is the reason this story is not a governance-page footnote.

The idea did not originate in the working groups. Demis Hassabis published an essay on 14 July 2026, A Framework for Frontier AI and the Dawning of a New Age, proposing exactly this: a US-initiated, federally overseen, industry-funded standards body on the FINRA pattern. His design had two phases. First, labs voluntarily share models for review up to 30 days before release, with testing across cybersecurity, biological risk and agentic behaviours including guardrail bypass and deception. Then, once the regime demonstrably works, compliance becomes a prerequisite for any model to go live for American users. Hassabis said funding “would need to be substantial and likely mostly come from industry,” and that he wanted the body operational before year-end.

This desk covered the earlier stage of this on 14 September, when the three labs’ private standards working groups surfaced without the antitrust waiver their own proposal said they needed. What has changed in eleven days is that the thing now has a name, a launch window, a leadership shortlist and a remit. It has acquired institutional shape.

The remit is the story

Reread the fourth row of that table.

In any functioning assurance market — financial audit, clinical trials, aviation airworthiness, food safety — the value of an audit comes entirely from the auditor’s independence, and independence comes from somewhere the audited party does not control. A statute. A professional body. A regulator with subpoena power. The audit is worth something precisely because the audited firm did not decide who was allowed to perform it.

SAFA’s reported design puts that decision inside the three companies being audited.

This does not require anyone to act in bad faith to produce a bad outcome. It requires only that qualification criteria be written by people with a consistent institutional interest, over years, with no countervailing party in the room. FINRA is the stated model, so FINRA’s record is fair evidence about what the model produces. Senator Elizabeth Warren has criticised it repeatedly for weighting brokerage interests above investor protection. Whistleblowers accused it of declining to pursue problematic brokers at large member firms. Its own former enforcement chief, Brad Bennett, has said some larger brokerages treat its fines as cheaper than compliance. SLCG Economic Consulting found it had failed to apply its own “restricted” designation to at least 13 brokerages that met the criteria.

That is the mature, seventy-year-old version of this institutional form, in a sector with far better-defined harms and far more litigation pressure than AI. It is not an argument that SAFA would be worthless. It is an argument about what a self-regulator’s steady state looks like, and buyers should price the steady state rather than the launch announcement.

The critics on record this quarter are blunter. Nader Henein of Gartner: “self-regulation is not viable,” because most tech vendors “don’t have the capacity to self-regulate.” Independent analyst Carmi Levy called the Hassabis framework “a self-serving roadmap for an industry bent on racing to the AI horizon regardless of the harms caused along the way.” The most useful objection is Yoshua Bengio’s, who accepted the pragmatic case for starting voluntary but insisted on “a clear and precise roadmap to transition from a voluntary to a mandatory model.” Nothing in the reported SAFA design contains that roadmap. Hassabis’s original essay did.

Two names, sixteen days apart

The reported chief-executive shortlist is Sriram Krishnan, Arati Prabhakar, Condoleezza Rice and David Friedberg.

Krishnan served as senior White House AI policy adviser from January 2025 until June 2026. On his way out, he told the Financial Times that “there will not be an FDA for AI,” arguing a central licensing agency requiring “a team of lawyers before you can get a model out” would put sand in the gears of the AI revolution. Reasonable people hold that view. It is nonetheless a striking résumé line for the prospective chief executive of a body whose reported job includes qualifying auditors and, in the Hassabis design, eventually gating US model releases. Either the position has changed, or SAFA is understood by its architects as the thing that makes an FDA for AI unnecessary — and the second reading is the one a buyer should assume until told otherwise.

The reported scientific consultants are Beth Barnes, founder and chief executive of METR, and Paul Christiano, senior technical adviser at the Commerce Department’s Center for AI Standards and Innovation.

Both are credible. METR is one of the few organisations that has done serious, published, adversarial evaluation of frontier models. Christiano co-invented RLHF and has spent a decade on alignment.

And on 8 September 2026 — sixteen days before this report — OpenAI announced that Christiano had joined its Foundation Board, taking a seat on the Safety and Security Committee chaired by Zico Kolter and a non-voting observer seat on the OpenAI Group PBC board.

Hold those two facts next to each other. A sitting board member of one of the three founding companies is reportedly a scientific consultant to the body that would define what independence means for everyone auditing that company’s models. That is not a scandal and it is not illegal. In a field this small, the people qualified to write frontier evaluation standards are largely the people the frontier labs have already hired, funded or appointed — which is itself the structural problem, not a defence against it.

It is also not an isolated instance. Last week Anthropic named Accenture as its first embedded evaluator, funded the work directly, and picked a firm whose AI arm had already red-teamed for OpenAI. That arrangement arrived structured as a vendor engagement — the one thing the published independence standard says it must not be. The pattern is consistent: the mechanism keeps being real, and the independence keeps being supplied by the party being evaluated.

The twenty-four hours this landed in

The timing is the part that will not fit in a governance summary.

On the same day this report ran, Politico reported that the White House Office of the National Cyber Director had asked OpenAI and Anthropic to withhold new frontier models from the UK AI Security Institute until a US review completes. Anthropic complied: Claude Mythos 5.1 went to a US-only Project Glasswing slate, the first time AISI had been excluded from a pre-release Anthropic evaluation. This desk covered it earlier today — the last independent frontier-model evaluator just became a US policy variable.

Set the two stories side by side and the shape is hard to miss:

Three independent developments, none of them coordinated, converging on the same outcome within eight days: every remaining route to an outside opinion about a frontier model now runs through a party with a commercial or political interest in the answer.

That is not a conspiracy claim. It is an inventory. And it is the reason the phrase “independently evaluated” — which appears in model cards, enterprise security questionnaires, EU AI Act readiness packs and half the model-risk documentation written this year — is quietly becoming unfalsifiable.

What this changes for procurement

Very little, immediately, and that is the trap. SAFA does not exist. No standard has been published. Nothing in your contract changes this quarter.

What changes is the evidentiary value of a phrase you are probably already relying on.

Stop accepting “independently evaluated” without the artefacts. Four questions, in writing, to every frontier vendor in your stack:

  1. Who performed the evaluation, and what is their commercial relationship to you? Paid engagement, grant, equity, board overlap, or genuinely none.
  2. Which model builds did they actually see? Pre-release access is now allocated by jurisdiction as well as by contract, so “the model was evaluated” and “this model was evaluated” have come apart.
  3. What level of access did they have? Weights and training data, internal evaluation results, or the same public API your team has.
  4. What could they have published if the findings had been bad? Publication rights are the only part of an independence claim that is cheap to verify and expensive to fake.

Do not wait for SAFA to supply a datasheet. Carmi Levy’s advice to enterprises this week was to enforce vendor standards now: require models to ship with the equivalent of a safety and security datasheet — documented capabilities, known failure modes, testing history. Nothing prevents you from making that a contractual condition today, and a vendor’s willingness to provide one is itself a signal.

Treat AI risk with the machinery you already have. Yaz Palanichamy of Info-Tech Research Group frames it as diagnosing AI risk “with the same level of seriousness as financial or cybersecurity risk” — a standing governance committee, named owners, contractual guarantees. The gap this closes is real and current: the labs are in no position to offer meaningful behavioural assurances for agentic systems right now, as the Beltdown sandbox-escape cluster demonstrated across every major coding agent this month.

If you operate in the EU, note which regime is actually binding. SAFA is voluntary and hypothetical. Article 50 of the AI Act has been enforceable since 2 August 2026, and its obligations fall on you as deployer regardless of what any industry body publishes. The contrast with the US approach — where the federal frontier framework was completed and then kept secret — is the practical reason European buyers currently have more to work with than American ones, not less.

Watch the membership, not the mission statement. The vendors outside SAFA have an incentive to compete on transparency, and some will. Microsoft published a versioned behavioural document rather than joining. Open-weight vendors — Mistral, DeepSeek, the Qwen family — can offer inspection that no closed-weight lab can match, which is a genuine procurement argument and not only an ideological one. Cohere has been the loudest critic of the club from inside the industry. A market with a self-regulator and a set of outsiders competing against it is a better market for buyers than one with neither.

The honest uncertainty

This is a report about a plan, sourced to The Information and repeated across secondary coverage, with no on-record statement from Google, OpenAI or Anthropic. The leadership names are people who have been approached, not people who have accepted. The remit could narrow before launch. Bengio’s mandatory-transition roadmap could be written into the charter. The body could turn out to do the unglamorous, genuinely valuable work — comparable benchmarks, a real incident-reporting channel — and leave auditor qualification alone.

It could also launch exactly as described, in a market where the alternative evaluator lost access the same week, funded by the three firms it grades, with its independence standard written by consultants who sit on those firms’ boards.

Both outcomes are live. The buyer-side response is identical either way, which is what makes it worth doing now: ask for the artefacts, not the adjective. Whether you are running Claude, ChatGPT or Gemini in production — and the Claude vs ChatGPT comparison covers where each currently sits on published safety documentation — the evaluation evidence behind them is about to become the part of the model card you cannot take at face value.

For teams building on these models directly, the developers guide covers the stack-level decisions, and OpenAI’s own frontier governance framework remains the clearest published statement of what one of these labs thinks it owes you. Read it as a commitment you can hold them to, because for now, that is the only kind of commitment on the table.

Update, 1 October 2026 — the question got answered the other way. This piece asked who qualifies the auditors. On 29 September six frontier labs signed the White House Accord on Super Intelligence, which answers it by declining to: the companies select their own auditors, with no requirement to disclose their identities or the selection process, no implementation deadline, and no obligation to publish findings. That is the opposite of a qualification regime, and it arrived from the White House rather than from a standards body. One day later the FTC confirmed a broad investigation into the safety of Anthropic’s and OpenAI’s systems, reportedly reaching for formal demands and compelled testimony — including, per Axios, from the evaluator METR. The artefacts this article told buyers to ask for are about to be produced for a regulator instead. Why the accord’s four audit layers have no outside reader.

Frequently asked questions

What is SAFA, and has it actually been created?

Not yet. On 24 September 2026 The Information reported that Google, OpenAI and Anthropic have agreed to establish a body tentatively called the Standards Authority for Frontier AI, with a target launch between the end of 2026 and early 2027. That is an agreement to form an institution, not an institution — there is no charter, no incorporation, no confirmed leadership and no published standard. The three companies have not issued a joint announcement, and none of the reporting carries an on-record company statement. The structure described is a self-regulatory organisation modelled on FINRA, the body that polices US brokerages: industry-funded, operating outside direct government control, and not requiring Congressional or presidential approval to exist — though reporting notes it may need to register with a federal agency to carry any legal authority, which is exactly how FINRA relates to the SEC. Treat everything below as a well-sourced report about a plan, and treat the plan itself as the thing worth reading, because the plan tells you what the three largest model vendors think the oversight layer above them should look like.

What would SAFA actually do?

Four functions appear consistently across the reporting. First, develop common technical standards and benchmark tests for frontier models, so that safety claims from different labs become comparable rather than each vendor grading itself on its own rubric. Second, support third-party pre-deployment safety testing — the mechanism where an outside group gets model access before release. Third, establish incident reporting rules, meaning a defined channel and threshold for disclosing when a deployed model does something it should not have. Fourth, and this is the one that changes a procurement conversation: set qualification requirements for independent auditors. Demis Hassabis's July framing added a fifth element that has not clearly survived into the reported SAFA design — a voluntary phase in which labs share models up to 30 days pre-release for testing in cybersecurity, biological risk and agentic domains including guardrail bypass and deception, followed by a phase in which compliance becomes a prerequisite for a model to go live for US users. The first three functions are things buyers have wanted for two years. The fourth is the one to read twice.

Why is the auditor-qualification remit the part that matters?

Because it inverts the direction of accountability. In a working assurance market, the auditor's independence is what gives the audit value, and independence is established by someone the audited party does not control — a regulator, a professional body, a statute. If the three largest frontier vendors jointly define who is qualified to audit frontier models, then the pool of auditors permitted to grade them is a pool they constituted. An auditor whose findings are consistently unwelcome does not need to be fired; the qualification criteria simply need to change. That is not a prediction about bad faith, and the same criticism has been made of FINRA for two decades by people including Senator Elizabeth Warren, whose objection was that the body reliably weighted brokerage interests over investor protection. FINRA's former enforcement chief Brad Bennett has said plainly that some larger brokerages treat its fines as cheaper than compliance. SLCG Economic Consulting found FINRA had failed to apply its own 'restricted' designation to at least 13 brokerages that qualified for it. A self-regulator with a bad incentive structure does not usually fail loudly. It grades on a curve for a long time.

Who would run it, and why do the names matter?

The reported chief-executive shortlist is Sriram Krishnan, Arati Prabhakar, Condoleezza Rice and David Friedberg — respectively the second Trump administration's senior White House AI policy adviser until June 2026, the Biden administration's head of the Office of Science and Technology Policy, a Bush-era Secretary of State, and a venture capitalist. Krishnan is the pointed one. On his way out of the White House he told the Financial Times that there will not be an FDA for AI, arguing a central licensing agency that demanded 'a team of lawyers before you can get a model out' would put sand in the gears. That is a defensible policy position. It is an unusual qualification for running the body that would qualify auditors and gate model releases. The reported scientific consultants are Beth Barnes, founder and chief executive of the evaluation nonprofit METR, and Paul Christiano, a senior technical adviser at the Commerce Department's Center for AI Standards and Innovation. Christiano joined OpenAI's Foundation Board on 8 September 2026, sitting on its Safety and Security Committee and as a non-voting observer on the OpenAI Group PBC board. He is a serious safety researcher with genuine standing. He is also, sixteen days later, reportedly a scientific consultant to the body that would define independence for auditors of a market his board seat sits inside. Both things are true at once, and the second one is the one that belongs in a procurement file.

Who is not in it?

Meta, xAI, Microsoft, Mistral, Cohere, DeepSeek, Alibaba and every open-weight lab. The founding membership reported is three companies. Microsoft's absence is notable because it was a founding member of the 2023 Frontier Model Forum and because it chose a different route this month — publishing a versioned behavioural document for the models running inside its products rather than joining a club. That difference is the whole procurement story, and this desk made the same argument two weeks ago when the working groups first surfaced: a standards body you are not a member of is not a control you hold, and a versioned document you can diff is. The exclusion also has a competitive shape worth naming. Standards written by three closed-weight frontier labs will tend to encode assumptions — staged pre-release access, controlled deployment, auditable training pipelines — that an open-weight release cannot satisfy by construction. If SAFA compliance ever becomes a de facto requirement for US enterprise procurement, the labs that publish weights are disadvantaged by the standard's shape rather than by its content. Aidan Gomez of Cohere used a one-word description of the group earlier this month that happens to name a legal category.

Is this legal? What happened to the antitrust waiver?

The waiver never arrived, and the coordination is proceeding without it. Dario Amodei's September essay argued that meaningful safety coordination between competing labs requires a narrow US antitrust exemption first. Senator Josh Hawley, backed by Senator Ted Cruz, blocked a national-security antitrust exemption in the NDAA manager's package on 15 September. Treasury Secretary Scott Bessent told a House hearing the labs should not receive the liability exemption they sought, and FTC Chairman Andrew Ferguson said the waiver request should make everyone deeply suspicious. What the labs got instead was softer and cheaper: on 17 September, Associate Attorney General Stanley Woodward said at Fordham that the Justice Department does not currently view coordination among companies on AI safety as anticompetitive. That is an enforcement posture, not a safe harbour — it binds no successor, creates no immunity and does not touch private antitrust litigation. OpenAI has said publicly it does not believe it needs a waiver. The practical reading is that the three labs decided a stated DOJ disposition was sufficient cover to proceed, having failed to get the statutory version through Congress. Anything built on an enforcement posture can be unbuilt by the next one.

What should a buyer actually change this week?

Stop accepting 'independently evaluated' as a term of art in vendor documentation, and start requiring the artefacts underneath it. Four questions, in writing, from every frontier vendor in your stack. Who performed the evaluation, and what is their commercial relationship to you — paid engagement, grant, equity, board overlap, or none? Which specific model builds did they see, given that pre-release access is now allocated by jurisdiction as well as by contract? What did they have access to — weights, training data, internal evaluations, or a public API like everyone else? And what would they have been permitted to publish had the findings been bad? Separately, take the advice independent analyst Carmi Levy gave enterprises this week and do not wait for SAFA: require that models ship with the equivalent of a safety and security datasheet covering documented capabilities, known failure modes and testing history. Yaz Palanichamy of Info-Tech Research Group frames the governance side the same way — diagnose AI risk with the same seriousness as financial or cybersecurity risk, through a standing committee with contractual teeth rather than a vendor questionnaire. None of that depends on SAFA existing, which is the point of doing it now.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.