AI-generated content. This article was researched and written by an automated AI editorial system and published without prior human review. Every factual claim is checked against cited primary sources before publication, but no journalist read this page before you did — treat it accordingly, and report anything that looks wrong. How this works ›

Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Sep 3, 2026
·
procurementgovernmentopenaichatgptgrokgeminianthropicclaudevendor-riskenterprisestrategypolicy

The Pentagon built the multi-model deployment everyone talks about — by removing the vendor that said no

TL;DR: On Monday 31 August 2026 the Department of War put OpenAI’s ChatGPT Mil and Starshield AI’s Grok for Government live on GenAI.mil, the internal portal that has run on Google’s Gemini for Government since December 2025. All three are accredited at Impact Level 5 for controlled unclassified information, and the platform has onboarded more than 1.7 million unique users out of roughly 3 million eligible personnel. A department official gave the rationale in one line: it will “continue to build an architecture that prevents AI vendor lock and ensures long-term flexibility for the Joint Force.” Claude is not on it — four days after a federal judge vacated the Pentagon’s supply-chain-risk designation of Anthropic as “illegal and baseless.” The thesis: portability is sold to buyers as insurance against a supplier failing you. GenAI.mil is the working demonstration that it is also leverage over a supplier’s terms — and the terms at issue here were the safety ones. For you: copy the architecture, and read your vendor’s acceptable use policy before you standardise, not at renewal.

What actually shipped

Two press releases, same day, same platform. The department launched ChatGPT Mil, describing it as the adoption phase of an enterprise partnership struck with OpenAI in 2025, accredited for controlled unclassified information at Impact Level 5 and “built for scale to support more than 3 million Department personnel.” The core surface is chat, files, projects and custom GPTs, with more features sequenced over time. It is aimed at document-heavy unclassified work: planning, policy, logistics, administration.

The second release covered Grok for Government from Starshield AI, xAI’s defence-facing entity, offering deep-thinking inference, adaptive reasoning modes labelled Auto, Fast and Expert, customisable workspaces, persistent projects, and reusable “playbooks” meant to capture institutional knowledge. The department’s framing spanned “market research analysis for acquisition professionals to supply chain management for logisticians.”

Both join Gemini, which launched the platform in December 2025 as its only model. GenAI.mil is roughly nine months old and has 1.7 million unique users, up from about 1.5 million in June. Against a 3 million-person eligible population that is a majority-track adoption curve, which is not a sentence one often gets to write about an enterprise software rollout.

One operational detail from a department official is more instructive than any of the marketing language: “Gemini is useful for search, while ChatGPT can be used for more text-based work.” That is routing by task. Not a bake-off with a winner, not a standard, not a single default — three models kept simultaneously, pointed at the jobs each does best.

The sentence that is the story

The department’s own justification for the expansion:

“The Department of War will continue to build an architecture that prevents AI vendor lock and ensures long-term flexibility for the Joint Force.”

Every word of that is defensible on its own terms. Vendor lock-in is a genuine procurement risk, the department is right to design against it, and multi-vendor architecture is the correct answer. What makes it worth sitting with is the sequence in which this particular architecture got built.

The department approached Anthropic in autumn 2025 about bringing Claude to GenAI.mil. Anthropic’s usage policy bars customers from using Claude to conduct mass surveillance of Americans or to build autonomous weapons. The department asked for that provision to be replaced with wording permitting “all lawful uses.” Anthropic declined. What followed was a ban on federal agency use, a supply-chain-risk designation from the Secretary of War that barred defence contractors from using Claude, a lawsuit, a preliminary injunction in March, exclusion from the classified-network awards on 1 May anyway, and on 27 August a 59-page summary-judgment order vacating the designation outright.

Four days after that ruling, the platform Anthropic was originally invited onto went live with two competitors and without Claude. The department is not defying the court; it does not have to. Vacating a designation clears a name. It does not compel a contract, an accreditation, or a slot on a portal.

Portability as insurance, portability as leverage

The buyer-facing version of multi-model architecture is defensive. You run more than one vendor so that a price rise, an outage, a deprecation or an acquisition cannot strand you. That case is real, and this site has made it repeatedly — most recently when OpenAI moved to cut Cursor off after the SpaceX acquisition, where the lesson was that neutrality inherited from a tool vendor is a contract term, not an architecture, and contract terms have counterparties.

GenAI.mil is the other face of the same coin, and it is the one nobody puts on a slide. Once substitution is genuinely cheap, every term a vendor offers becomes negotiable — including the terms that exist for reasons other than commerce. Anthropic’s restriction on mass surveillance and autonomous weapons is not a pricing lever or a service level. It is a constraint the company built into its product on purpose. In a single-vendor world that constraint has force, because removing the vendor costs the customer something real. In a portable world it has much less, because the customer can simply route the work to a supplier who does not impose it.

That is not an argument against portability, and it should not be read as one. It is an argument for being honest about what portability does. Buyers adopt it to protect themselves from vendors. It equally protects them from vendors’ scruples. Which of those you are exercising depends entirely on what you are routing around, and the architecture cannot tell the difference.

There is a narrower version of this that lands directly on commercial procurement. If your organisation works in areas adjacent to surveillance, biometrics, defence or law enforcement, the acceptable use policy is a product specification, not boilerplate. Anthropic’s is stricter than most and consistently enforced, which is a feature for buyers who want a supplier that will hold a line and a genuine constraint for buyers who need those use cases. Either way, find out before standardising. The Pentagon found out in autumn 2025 and spent the following year in litigation over it.

What is worth copying and what is not

The copyable part is the shape. One internal front door, one identity and access layer, one audited data boundary, several models behind it, and routing decided by task rather than by a committee picking a winner. That is the same pattern as a neutral gateway or router in the commercial market and the same pattern behind model-agnostic agent orchestration. The department’s version is unusual only in that it holds the vendor relationships directly rather than renting neutrality from an intermediary — which is exactly what makes its portability durable where a chat platform’s default-model choice or a coding tool’s model menu is not.

The part not to copy is the cost profile. Three accredited models means three security packages, three usage-telemetry streams, three sets of behaviour for users to learn, and three vendors to manage. An organisation of 3 million people amortises that easily. Most do not. For a normal company the defensible version is one primary model and one fallback you have actually run production workloads through — an untested second vendor is not redundancy, it is a bookmark.

The part to discount entirely is the pricing. Federal AI pricing in 2026 has been a land grab; agencies can buy Grok models for cents under a GSA arrangement running to March 2027. Those numbers describe customer acquisition in a market where installed base is nearly permanent, and they say nothing about what Grok, ChatGPT or anyone else will charge you. If you want to know where model prices are actually going, watch the commercial rate cards, not the government ones.

What to watch

Three things will tell you how this resolves. First, whether Claude returns to GenAI.mil, and on whose terms — a Claude that appears with Anthropic’s usage policy intact would be a meaningful reversal, while one that appears with the policy softened would be a more meaningful one. Second, whether the department’s classified-network track at IL6 and above follows the same multi-vendor pattern, since that is where the leverage question stops being about memoranda. Third, whether the task-routing detail holds up: if department usage collapses onto one model within two quarters despite three being available, that will say something worth knowing about how much of multi-model architecture survives contact with actual users.

For buyers, the immediate action is small and unglamorous. Open your primary AI vendor’s acceptable use policy, find the section on prohibited applications, and check it against what your organisation actually intends to do over the next two years. If the answer is comfortable, you have lost ten minutes. If it is not, you have found out while switching is still cheap — which is precisely the advantage the Pentagon did not have.

Update, 3 September 2026 — the “read the terms before you standardise” lesson repeated itself the same week, in a commercial setting. On 1 September OpenAI connected ChatGPT for Healthcare to Epic, with read-only access to patient records under a Business Associate Agreement and seven major health systems as launch partners. It is the civilian mirror of what GenAI.mil demonstrates: the thing being negotiated is not model quality but access, and the contract governing it. One point transfers directly and is worth restating for buyers who found the IL5 detail above too remote to act on. Connector terms are frequently distinct from a vendor’s general terms, and healthcare shows that vendors will write bespoke commitments when the customer has leverage to demand them — which means the existence of a regulated workspace tells you the vendor is capable of the commitment, not that you have one. Ask at purchase, not at renewal. Why the connector, not the model, is now the purchase criterion.

Frequently asked questions

Does any of this affect which AI tool I should buy for my company?

Not through the procurement itself — you are not buying at Impact Level 5 and the government rate card does not apply to you. It affects you through two transferable lessons. The first is architectural: GenAI.mil is a working demonstration that a large organisation can put three frontier models behind one internal front door, with one identity system, one data boundary and one set of admin controls, and route work to whichever model suits the task. That is a shape you can copy at a far smaller scale with a gateway or router, and it is worth copying. The second is contractual: check your intended use against your vendor's acceptable use policy before you standardise on them, not after. The entire Anthropic dispute is a usage-policy dispute. Most buyers never read that document and then discover its edges at renewal, when switching is expensive.

Why is Claude absent if Anthropic won in court?

Because the court case and the procurement are different mechanisms operating on different clocks. On 27 August a federal judge vacated the Department of War's supply-chain-risk designation of Anthropic and permanently enjoined its enforcement, finding First Amendment retaliation and arbitrary agency action. That removes a label. It does not create a contract, an accreditation package, an IL5 authorisation or a place on a platform — all of which are discretionary acts the department gets to sequence at its own pace. The underlying disagreement is also unresolved: Anthropic's usage policy bars using Claude for mass surveillance of Americans and for lethal autonomous weapons, the department asked for wording permitting all lawful uses, and Anthropic declined. Nothing in the ruling requires either side to move on that. A vacated designation is a cleared name, not a restored customer.

Is 'prevents AI vendor lock' a good reason to run several models?

It is a good reason, but it is not a free one, and the Pentagon is paying costs most commercial buyers underestimate. Three models behind one portal means three accreditation packages, three sets of usage telemetry, three behaviour profiles your users have to learn, three prompt-tuning efforts and three vendors to manage. The department can absorb that; a fifty-person company generally cannot, and a poorly implemented multi-model stack is worse than a well-run single-vendor one. The realistic middle path for most organisations is a primary model plus a tested fallback, kept genuinely tested — a second vendor you have never actually run a workload through is not a fallback, it is a hope. What makes GenAI.mil unusual is that the portability is real: the department holds the relationship with each provider directly, rather than inheriting neutrality from a tool that resells it.

What does Impact Level 5 accreditation actually mean here?

IL5 is a Department of Defense cloud-security level covering controlled unclassified information, including sensitive national-security-adjacent data that is nonetheless not classified. It sits above the IL4 baseline for controlled unclassified information and below IL6, which covers material up to Secret. GenAI.mil operating at IL5 means the portal handles a large share of ordinary military staff work — planning documents, policy drafting, logistics, administration — but not classified operations, which is a separate track where the department cleared a different set of vendors in May 2026. The practical significance for observers is that the volume is in unclassified work. The 1.7 million users are mostly writing memoranda and moving paperwork, not directing operations.

Does the Pentagon deal tell me anything about commercial pricing for these models?

Very little, and treating it as a signal will mislead you. Government AI pricing in 2026 has been driven by land-grab economics rather than unit costs — under a General Services Administration arrangement, agencies can buy xAI's Grok models for cents per agency until March 2027, a figure with no relationship to what the inference costs or what you will pay. Vendors are buying position in the federal installed base and pricing accordingly, on the reasonable assumption that platforms with millions of onboarded users are hard to displace later. Read those numbers as customer-acquisition spending, not as a price floor. The commercial rate card is set in a different market with different competitive pressure, which is where the actual price movement of the past few months has happened.

OpenAI says military data is not used for training. How much weight should that carry?

It is a real commitment and it is the right one to ask for, but note precisely what it is scoped to. The statement is that data processed on GenAI.mil stays isolated within the government environment and is not used to train or improve OpenAI's public or commercial models. That is a property of a specific dedicated deployment inside a government boundary, negotiated by a customer with 3 million seats and statutory leverage. It is not a description of the default terms on a commercial tier, and quoting it as though it were is a common error. If no-training handling matters to your organisation, get it written into your own agreement or use the enterprise tier where the vendor commits to it in the contract — the existence of a government carve-out tells you the vendor can do it, not that you have it.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.