AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Sep 10, 2026
·
policysecuritychinaopen-weightsapi

A federal advisory just told US model providers to serve degraded models to suspected distillers — and to vary it so quality evaluation cannot detect it

TL;DR: On 8 September the NSA, CISA, and the FBI issued joint advisory AA26-251A, naming DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI for industrial-scale distillation of US frontier models — billions of tokens across millions of requests since late 2024, against Claude, GPT, Gemini, and Grok. The accusation is the story everyone ran. The story for buyers is mitigation two, which advises US providers to subtly alter responses to suspected accounts — differential privacy or serving degraded models — and to vary it across requests specifically to complicate quality evaluations. Then read the detection indicators: sustained usage around the clock without human variation, new subscriptions immediately running at maximum usage, shared accounts across multiple IPs and user agents, traffic via aggregators that obfuscate metadata. That is a description of a normal production agent fleet. The advisory recommends no notification and no appeal path, and offers zero guidance to downstream buyers. Nobody has shown this happening to a legitimate customer. But the countermeasure is now officially endorsed, and its defining feature is that your benchmarks cannot see it.

Update (10 September 2026): two days after being named in this advisory, DeepSeek shipped V4.1-Flash with MIT-licensed weights and announced that deepseek-v4-pro traffic routes to it on 14 September. That routing change is a useful illustration of the canary-set argument below: the model behind a stable identifier changes, nothing errors, and the only signal is a benchmark regression on reasoning tasks.

What the advisory alleges

The factual core, briefly, because it is the least surprising part.

CompanyAlleged extractionUsed for
DeepSeekMultiple versions of Claude, GPT, GeminiTraining data for R1 and V3
Moonshot AIClaude FableKimi K3
Moonshot AIGPT-4o outputKimi K2
AlibabaUS frontier modelsImproving the Qwen family
MiniMax, StepFun, Z.AINamed in the same campaignNot individually itemized

The agencies say the activity has run since late 2024, that it moved billions of tokens across millions of requests, and that the Chinese government is likely aware. Methods: fraudulent accounts, bulk premium subscription harvesting, native APIs, remote cloud providers, and gray-market proxy resellers the advisory calls transfer stations — services that resell frontier access at a discount while bypassing geographic restrictions and obscuring who is asking.

For readers following this thread, none of it is new in kind. It escalates Anthropic’s June accusation against Alibaba, the Claude Code tracking detail that followed in July, and the White House’s Moonshot/Kimi K3 claim and sanctions threat. What is new is the institutional weight: three agencies, one document, six named companies.

The recommendation nobody led with

The advisory gives US providers three immediate defensive approaches. One and three are unremarkable — comprehensive detection of anomalous prompts and accounts, and cross-organization intelligence sharing across model providers, cloud platforms, and API aggregators.

Mitigation two is the one to read twice:

Deploy targeted response changes that subtly alter responses to suspected malicious distillation attempts — such as differential privacy or serving degraded modelsvaried across requests to complicate quality evaluations.

Take that clause apart. The recommendation is not to block. Blocking is visible: you get an error, you file a ticket, you find out. The recommendation is to keep serving while making the output worse, and to randomize the degradation so that measuring quality does not reveal it.

That is not an incidental property. It is the stated design goal. The countermeasure is specified to defeat exactly the method — run a benchmark, compare scores — that any competent buyer would use to check whether they are getting what they paid for.

Against an actual distillation operation, this is clever policy. Poisoned training data is a far better outcome than a blocked account that simply re-registers, and it raises the adversary’s cost in the right currency. The logic is sound.

The problem is the aperture.

The signature describes your agent fleet

Set the mitigation beside the advisory’s own detection indicators:

Four of five describe ordinary, legitimate, high-volume production usage. The fifth describes the architecture the market has been converging on all year.

Be precise about the claim here, because the distinction matters: there is no evidence any US provider is degrading legitimate customers, none has announced such a practice, and providers have obvious commercial reasons not to. The advisory recommends; it does not report. Real distillation operations look different in aggregate than a startup’s backfill, and provider fraud teams are not naive.

But three things are now true at once. A federal advisory has formally endorsed silent quality degradation as a countermeasure. The published signature for triggering it includes behaviors that are indistinguishable from normal automation. And the document specifies no notification requirement, no appeal path, and no guidance whatsoever for the downstream buyers who would absorb a false positive. That combination is what changes, and it changes something specific: model quality moved from assumed to unverified.

Why your existing evaluation will not catch it

Most teams check model quality one of two ways, and the advisory’s phrasing defeats both.

Public benchmarks were already weak evidence for your workload, and they are weaker here — anything public is plausibly special-cased, which is precisely why independent evaluators moved to private test sets when Artificial Analysis rebuilt its index. A degradation applied per-account does not show up in anyone’s public leaderboard, because the leaderboard is not running on your account.

Aggregate internal metrics — average quality score, thumbs-up rate, eval pass rate over a week — are exactly what varied across requests is designed to survive. Randomized per-request degradation shows up as slightly wider variance, not as a step change, and variance is the one thing nobody alerts on.

What actually works is dull and cheap:

  1. A private canary set. 30–50 fixed prompts with known-good reference outputs, scored automatically, run on a schedule. Private, versioned, never published.
  2. Two paths, same model. Run the identical set through a second route — different key, different region, or a gateway — and diff the scores. A delta between two paths hitting the same nominal model is a real signal; drift in a single path is not, because the model changes underneath you anyway.
  3. Alert on variance, not just the mean. The specified countermeasure widens the distribution before it moves the average.

None of this is exotic. Most teams running agents in production want it for ordinary regression reasons, and the outage-correlation lesson from 3 September already argued for two live paths. This advisory just added a second reason to build it.

What it means for Chinese open weights

A narrower question, worth separating cleanly from the above.

Running Qwen, GLM, or a DeepSeek model on your own hardware is not what the advisory targets. It addresses extraction from US APIs; it says nothing about deploying the resulting weights, and self-hosting open weights is neither alleged nor prohibited anywhere in the document.

What shifts is provenance risk. Three agencies naming six vendors, on top of July’s White House accusations, is a documented federal position — and it raises the odds of downstream sanctions or Entity List action touching a model line somebody has already built on. That is a real consideration for anyone whose contracts or export posture care where a model came from, and it applies to the open-weight price floor that GLM 5.3 Flash and Qwen Flash Next established and to the capability questions GLM 5.3 raised in August.

The proportionate response is not to rip out working deployments. It is to know which ones you could replace, how fast, and at what cost — and to keep the abstraction layer that makes the answer “a config change” rather than “a quarter.”

What to do about it

  1. Build the canary set this month. Private prompts, reference outputs, scheduled runs, variance alerting. It is a day of work and it is the only thing that answers the question.
  2. Keep two live paths to your critical model. Not for failover alone — for the comparison. A quality delta between paths is your detection mechanism.
  3. Make your account obviously legitimate. Real business entity, accurate contact details, seat count that bears some relationship to volume, and a heads-up to your provider before a large backfill or load test. The signature that flags you is behavioral; the context that clears you is administrative.
  4. Ask your provider directly, in writing. Do you implement response degradation for accounts flagged as suspicious, and will you notify us if our account is flagged? It is a fair procurement question with a documented federal advisory behind it, and the answer belongs in your contract.
  5. Treat gateways as measurement and portability, not anonymity. The advisory flags aggregators and recommends cross-provider intelligence sharing. Route through one because it lets you move and compare — and pick tooling on that basis, whether you are choosing coding tools or comparing Claude against ChatGPT.
  6. Separate the geopolitics from the deployment decision. Named-vendor risk is a procurement and provenance question. It is not a reason to abandon open weights, and conflating the two produces worse decisions than either concern alone.

The bottom line

The accusation in AA26-251A is the part that will be argued about, and it is largely a continuation of a fight that has been running since June. The durable change is one sentence of recommended mitigation.

A US federal advisory has now told model providers that a reasonable response to a suspected bad actor is to keep taking their money and quietly serve them a worse model, randomized so they cannot measure it — with no notification, no appeal, and no word to the buyers who might be caught by a signature that reads like ordinary automation.

Almost certainly this will never touch you. But “almost certainly” is a posture, not a measurement, and until last week nobody official had suggested that the quality of a paid API response was a variable providers should deliberately make unverifiable.

Quality you cannot measure is quality you are taking on trust. Build the canary set.

Frequently asked questions

What does advisory AA26-251A actually say?

It is a joint cybersecurity advisory from the NSA, CISA, and the FBI, issued 8 September 2026 and reported widely on 9 September. It alleges that six China-based AI companies — DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI — have run industrial-scale distillation campaigns against US frontier models since late 2024, pulling billions of tokens across millions of requests from Anthropic's Claude, OpenAI's GPT, Google's Gemini, and xAI's Grok. Specific claims include DeepSeek drawing on multiple versions of Claude, GPT, and Gemini to generate training data for its R1 and V3 models, and Moonshot extracting from Claude Fable for Kimi K3 and GPT-4o output for Kimi K2. The advisory says the Chinese government is likely aware. Methods described include fraudulent accounts, bulk premium subscriptions, and gray-market proxy resellers called transfer stations that bypass regional restrictions. The advisory is directed at US model providers; it contains no guidance for downstream buyers.

Why does the degraded-response recommendation matter to me if I am not distilling anything?

Because the recommendation is engineered to be undetectable and the detection signature is broad. Mitigation two advises providers to deploy targeted response changes that subtly alter responses to suspected malicious distillation attempts — differential privacy or serving degraded models — and to vary this across requests specifically to complicate quality evaluations. Now read the published indicators: shared accounts used from multiple IP addresses and user agents, sustained usage around the clock without human variation, anomalous subscription-to-usage ratios, and new subscriptions immediately running at maximum usage. A CI pipeline hitting an API from a rotating IP pool matches indicator one. Any always-on agent fleet matches indicator two by definition — that is what automation is. A team that buys seats and immediately runs a backfill matches three and four. None of that is malicious. All of it is on the list.

Is there any evidence providers are actually degrading legitimate customers?

No, and that distinction should be held firmly. The advisory is a recommendation, not a report of practice. No US provider has announced that it serves degraded models to suspected accounts, none has been shown to be doing it to a legitimate customer, and the recommendation is aimed at genuine adversarial extraction rather than at ordinary enterprise traffic. Providers also have strong commercial reasons not to silently degrade paying customers. The problem is narrower and still real: a federal advisory has now formally recommended a countermeasure whose defining property is that it defeats quality measurement, and it did so without recommending any notification, appeal, or disclosure path for accounts caught by mistake. That converts a previously theoretical concern into a documented, officially endorsed one. The right posture is instrumentation, not alarm.

How would I even detect this if it happened to my account?

Standard benchmarking will not do it, because varying degradation across requests is designed precisely to defeat aggregate quality evaluation. What works is a canary set: 30 to 50 fixed, private prompts with known-good reference outputs, run on a schedule, scored automatically, tracked over time. Two design points matter. First, keep them private — anything public is inside training data and inside whatever a provider might special-case, the same reason independent evaluators moved to private test sets. Second, run the identical canary set through a second path: a different API key, a different region, or a gateway, and compare. A quality delta between two paths hitting the same nominal model is a far stronger signal than a drift in your own scores over time, because it controls for the model being updated underneath you. This is straightforward instrumentation and most teams running agents in production should already have it for regression reasons.

Does this make Chinese open-weight models riskier to deploy?

It changes the legal and procurement risk more than the technical risk. Downloading and self-hosting Qwen, GLM, or a DeepSeek model is not distillation and is not what the advisory targets — it addresses extraction from US APIs, not use of the resulting weights. But an advisory naming six vendors, from three agencies, following earlier White House accusations against Moonshot and Anthropic's own claims against Alibaba, establishes a documented federal position. That matters for anyone whose procurement, customer contracts, or export posture is sensitive to provenance, and it raises the odds of future sanctions or Entity List action affecting a model line you have already built on. The practical response is not to rip out working open-weight deployments but to know which ones you could replace and how fast. Keep the abstraction layer, keep a tested fallback, and do not build anything load-bearing on a single named vendor's hosted endpoint.

What about routing through OpenRouter or another aggregator — does that help or hurt?

It cuts both ways and you should know which way for your setup. The advisory explicitly flags third-party aggregators that obfuscate user metadata as an access pathway of concern, and recommends cross-organization intelligence sharing across model providers, cloud platforms, and API aggregators. So routing through a gateway does not make you invisible; it arguably places your traffic in a category the advisory asks providers to scrutinize harder, and your identity is now legible to a sharing arrangement rather than to one vendor. The benefit runs the other way: a gateway is the cheapest place to run the two-path comparison described above, and it is the mechanism that lets you move a workload if one provider's quality drifts. Net, the routing layer remains the right architecture — but treat it as a portability and measurement tool, not as anonymity, and make sure your account is clearly attributable to a real business.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.