AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Sep 25, 2026
·
policygovernanceanthropicopenaieu-ai-actsafetyprocurementevaluations

The last independent frontier-model evaluator just became a US policy variable

TL;DR: Politico reported on 24 September 2026 that the White House Office of the National Cyber Director asked OpenAI and Anthropic to withhold new frontier models from the UK AI Security Institute until a US review completes. The stated rationale from a senior administration official: “Because they’re American companies and this has been our policy with every new frontier model that comes out.” Anthropic has already complied — Claude Mythos 5.1 (released 1 September) went to a US-only Project Glasswing slate, the first time AISI has been excluded from a pre-release Anthropic evaluation, with no timeline for restoring access. AISI director Henry de Zoete confirmed the institute retains pre-release access to some models, naming GPT-6 Astra. Why buyers should care: AISI is the only well-resourced evaluator in the stack that is neither vendor-funded nor vendor-selected and has pre-release access. Its coverage is now set by jurisdiction rather than capability — so “independently evaluated” is a claim about a specific model version, not about a vendor.

What happened

On 24 September 2026, Politico reported that the Office of the National Cyber Director — the White House unit that owns federal cyber strategy — asked OpenAI and Anthropic to hold new frontier models back from the UK’s AI Security Institute until the US government finishes its own review.

The request carries no legal force. There is no US statute compelling a private developer to route pre-release access through Washington, and the June 2026 executive order on frontier model review went out of its way to prohibit mandatory licensing, preclearance and permitting. This is a request made to companies that have strong commercial reasons to grant it.

A senior administration official’s rationale, as reported, was short: “Because they’re American companies and this has been our policy with every new frontier model that comes out.” The official framing is that US systems should be verified secure before models go to any partner.

Anthropic has already complied. Claude Mythos 5.1, released on 1 September alongside Fable 5.1, went to a US-only slate of organisations under Project Glasswing — the consortium Anthropic built for restricted-access cyber work. It is, on AI Weekly’s reporting, the first time AISI has been left out of a pre-release Anthropic evaluation. Anthropic said it would expand access to a broader set of domestic and international partners “as quickly as possible,” without a date.

OpenAI did not comment. AISI director Henry de Zoete confirmed the institute still holds pre-release access to some of the world’s most capable models, naming GPT-6 Astra.

So the picture is partial and undeclared: one frontier model confirmed withheld, one confirmed still in scope, and no published rule for where future releases land.

The part that actually reaches buyers

This reads like a geopolitics story. For anyone procuring AI tools it is an evidence story, and to see why it helps to lay out where frontier-model safety evidence actually comes from.

SourceIndependent of vendor?Pre-release access?
Vendor system cards and model cardsNo — self-reportedYes
Vendor-commissioned third-party red teamsFunded and selected by the vendorYes
Academic and non-profit indices (e.g. FLI)YesNo — works from public information
State-backed institutes (UK AISI and peers)YesYes

Only the bottom row has both properties. That is the whole point of a national safety institute, and it is why AISI’s evaluations carry weight in enterprise risk files that would not accept a vendor’s own testing.

The two middle rows have known problems this publication has documented. Vendor-commissioned evaluation has a structural funding issue that surfaced plainly when Anthropic’s Accenture embedded-evaluator arrangement and OpenAI’s red-team panel were announced within days of each other — the evaluator is paid and chosen by the evaluated, which is not disqualifying but is not independence either. Public indices like the FLI AI Safety Index are genuinely independent but cannot test what has not shipped.

Narrowing the fourth row by jurisdiction does not make any model less safe. It makes safety claims less externally checkable, which is a different problem, and it lands squarely on whoever in your organisation signs the model risk assessment.

The template was set in June

None of this is unprecedented. Pre-release access has been a controlled resource in US policy for more than a year, and the mechanism was demonstrated in public in June 2026, when Commerce cleared Mythos 5 for roughly 100 trusted US partners scoped to cybersecurity work while Anthropic’s public model stayed offline entirely.

What that episode established, and what this week extends, is a tiering of who gets to see a frontier model and when — a structure we mapped when capability tiers turned model access into a clearance rather than a purchase. The new element is that the tiering now has a nationality filter applied to evaluators, not just to customers.

It also runs alongside a transparency gap. The administration completed its frontier-model framework by the August deadline and has not published it. A review process whose contents are unpublished is difficult to cite as a substitute for one whose outputs are.

Why now

Two threads converged. The capability evidence has been getting harder to ignore — reporting around the request references tests in which systems were able to break into external targets including an Australian government portal, and both companies have briefed the UN Security Council on rising risk. OpenAI itself paused a frontier reinforcement-learning run over cyber-monitoring overhead in August.

Those results are exactly what makes a government want first look. They are also exactly what makes independent verification most valuable. Both statements are true at once, and this week the first one won.

The other thing that happened that day

The same twenty-four hours carried a second report on the same subject from the opposite direction. The Information reported on 24 September that Google, OpenAI and Anthropic have agreed to establish the Standards Authority for Frontier AI — a FINRA-style, industry-funded self-regulator targeting launch between end-2026 and early 2027, whose reported remit includes setting the qualification requirements for independent auditors.

Put the two side by side. The best-resourced evaluator that is neither vendor-funded nor vendor-selected lost pre-release access by government request. On the same day, the three vendors agreed to form the body that would decide who else qualifies to evaluate. Neither development caused the other, and that is rather the point — they converge without coordination. Full analysis of SAFA and what it does to the phrase ‘independently evaluated’.

What to do about it

Stop treating “independently evaluated” as a vendor property. It is a property of a model version, an evaluator and a date. Audit your AI risk documentation for sentences asserting third-party or institute evaluation, and pin each to a named version. Anything that cannot survive that should come out of the file rather than sit there as a stale assurance.

Add one line to your vendor questionnaire. Ask which external bodies received pre-release access to the specific version being supplied. It costs nothing and it is the only reliable way to find out, because vendors are not currently obliged to disclose an absence.

If you are in UK or EU public-sector procurement, ask in writing. Institute review is frequently treated as an implicit control in these processes. For at least one Anthropic frontier release, it is now absent.

Do not switch vendors over this. The models are unchanged; only who looked at them first has moved. There is no jurisdiction-clean alternative — the open-weight frontier carries its own evaluation gaps and its own history of licence changes mid-flight. If you want a practical hedge, it is the dull one: put the marginal hour into your own evaluation harness rather than waiting on someone else’s report. That was sound advice for developers before this week, and it is marginally better advice now.

For where the two vendors currently sit on capability and access, our Claude review and ChatGPT review track the shipping product rather than the policy layer — and for most buyers the shipping product is still the decision that matters. The difference is that the paperwork behind it now needs a version number.

Frequently asked questions

What exactly was reported, and by whom?

Politico reported on 24 September 2026 that the White House Office of the National Cyber Director asked OpenAI and Anthropic to withhold new frontier models from the UK AI Security Institute until the US government completes its own security review. The request is not a legal instruction — there is no statutory authority compelling a private company to route pre-release access through Washington — and it sits on top of the voluntary framework established by the June 2026 executive order, which explicitly prohibited mandatory licensing or preclearance. A senior administration official gave the rationale as: 'Because they're American companies and this has been our policy with every new frontier model that comes out.' The White House framing is that US systems should be verified secure before models are shared with any partner, foreign institutes included. Neither company issued a substantive public response at the time of the report; OpenAI declined to comment.

Has anything actually changed, or is this just a request?

It has already taken effect at one vendor. Claude Mythos 5.1, released on 1 September 2026, went to a US-only slate of organisations under Project Glasswing rather than to the usual pre-release evaluation set. AI Weekly's reporting describes this as the first time AISI has been left out of a pre-release Anthropic evaluation — a meaningful break, because AISI had evaluated the predecessor. Anthropic said it intends to 'expand access to a broader set of domestic and international partners as quickly as possible' but attached no timeline. On the OpenAI side, AISI director Henry de Zoete confirmed the institute still holds pre-release access to some of the most capable models, naming GPT-6 Astra specifically. So the current state is partial: one frontier model confirmed withheld, one confirmed still in scope, and no published rule for which future releases fall on which side.

Why does this matter to someone buying AI tools rather than writing AI policy?

Because it changes the quality of the evidence available when you do vendor due diligence, and most buyers have not noticed how thin that evidence base already was. There are broadly four sources of frontier-model safety evidence: the vendor's own system card, vendor-commissioned third-party red teams, academic and non-profit indices, and state-backed institutes with pre-release access. The first is self-reported. The second has a structural funding problem we covered when Anthropic's Accenture arrangement and OpenAI's red-team panel were announced — the evaluator is paid and selected by the evaluated. The third, including the Future of Life Institute's safety index, works largely from public information and cannot test a model before release. The fourth is the only category with both independence and pre-release access, and AISI is the best-resourced example of it. Narrowing that category by jurisdiction does not make models less safe. It makes their safety less externally checkable, which is a different problem and one that lands on whoever has to write the risk assessment.

Does this affect EU AI Act compliance obligations?

Not directly, but it affects how easily you evidence them. For general-purpose AI models with systemic risk, the Act places evaluation and adversarial-testing duties on the model provider, not on you as a deployer or downstream integrator. Nothing in the Politico report changes that allocation. Where it bites is practical: enterprise AI risk assessments and vendor questionnaires very commonly cite independent third-party evaluation as a control, and a growing number name AISI or an equivalent institute specifically. If your documentation contains a sentence to the effect of 'this model family is subject to independent pre-release evaluation by a national safety institute,' that sentence is now model-specific rather than family-specific, and for at least one Anthropic release it is no longer true. The fix is not complicated — name the model version and the evaluator, and re-check at each upgrade — but it does require someone to actually do it. Our [Article 50 explainer](/news/eu-ai-act-article-50-live-what-changes-for-ai-tool-users-2026-08-02/) covers the transparency side of the same regime.

Should I avoid US frontier models because of this?

No, and any vendor pitching that framing is selling you something. The models in question are the same models they were a week ago; what changed is who got to look at them first. There is also no clean alternative — the open-weight Chinese frontier has its own evaluation gaps, its own jurisdictional exposure, and a documented history of mid-flight licence changes. The reasonable response is narrower and duller. First, stop treating 'independently evaluated' as a stable property of a vendor and start treating it as a property of a specific model version with a specific evaluator and date. Second, if you are in UK or EU public-sector procurement, ask the vendor directly whether the version you are buying was reviewed by AISI or an equivalent, and get the answer in writing rather than inferring it from the vendor's general safety posture. Third, weight your own evaluations higher, because the external ones just got less predictable.

What is the background that made this happen?

Two threads converged. The first is the US frontier-review posture that has been building since the June 2026 executive order, which set up a voluntary 30-day federal review window while explicitly ruling out mandatory licensing, and which was followed by a completed frontier-model framework the administration has declined to publish. Pre-release access has been a controlled resource in US policy for over a year — the June 2026 Mythos 5 episode, where Commerce cleared roughly 100 trusted partners for a cyber-scoped release while a public model stayed offline, established the template. The second thread is capability evidence. Reporting around the request references tests in which systems were able to break into external targets, including an Australian government portal, and both companies have briefed the UN Security Council on rising capability. Those are exactly the results that make a government want first look. The tension is that they are also exactly the results that make independent verification most valuable.

What should I do about it this week?

Three things, none of them urgent but all of them cheap. Audit your AI vendor documentation for claims about third-party or institute evaluation and pin each one to a named model version and date — if a claim cannot survive that, delete it rather than leaving a stale assurance in a risk file. Add a question to your standard vendor questionnaire asking which external bodies received pre-release access to the specific version being supplied. And if you rely on a frontier model for anything where an unexpected agentic behaviour would be expensive, invest the marginal hour in your own evaluation harness rather than waiting for someone else's report; that was good advice before this week and it is slightly better advice now.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.