Meta shipped its best model yet with the weights still missing — and a price sheet that says what your data is worth
TL;DR: Meta released Muse Spark 1.3 on 2 September 2026, four weeks after 1.2. By the numbers it is the company’s strongest model — 75.4 on DeepSWE v1.1, 88.8 on Terminal-Bench 2.1, 98.1% MRCR retrieval in the 512k–1M band, and sixth of 636 on the Artificial Analysis Intelligence Index. Chief AI Officer Alexandr Wang told Axios it is “very competitive with frontier models.” The weights are not published. Neither are Muse Spark 1.2’s, pledged on 10 August to arrive “in the coming weeks.” What Meta did ship is a two-tier endpoint: standard at $1.25/$4.25 per million tokens with your data private, contributor at $0.10/$0.20 where “prompts and outputs may be used to improve Meta’s products.” The thesis: the gap between those two columns is not a discount. It is Meta’s own posted valuation of your traffic — roughly $1.63 per million tokens — and it is the most honest number the company published this week.
What shipped
Muse Spark 1.3 went live on Tuesday 2 September into two places at once: the Meta Model API, and Muse Code, Meta’s coding agent for the terminal and CI. It takes text, image and video input across a 1,048,576-token context window.
The headline results, from Meta’s own table:
| Benchmark | Muse Spark 1.3 | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|
| DeepSWE v1.1 (agentic SWE) | 75.4 | 74.0 | — |
| Terminal-Bench 2.1 | 88.8 | 86.7 | — |
| SWEAtlas CodeBase QnA | 59.4 | 52.7 | 53.5 |
| MRCR 256k–512k | 98.5 | — | 91.5 |
| MRCR 512k–1M | 98.1 | — | 73.8 |
| OSWorld 2.0 | 66.9 | — | — |
| DeepSearchQA | 89.4 | — | 93.0 |
| Agentic IF Index | 57.8 | — | 60.5 |
| GDPVal-AA v2 | 1754 | 1824 | 1710 |
Meta is not sweeping the board and does not claim to. It loses on deep search, on agentic instruction-following, and on the GDPVal economic-value composite. It wins on the coding cluster and it wins enormously on long context.
The independent corroboration is Artificial Analysis, which places 1.3 sixth of 636 models on its Intelligence Index. That does not validate any individual score in the table above, but it does confirm the general claim: Meta is back in the frontier conversation, which is not something we could write about Muse Spark 1.1 in July.
The number that reverses our last verdict
When we covered Muse Spark 1.1, the sharpest criticism we made was about long context. Meta advertised a million-token window and then scored well below GPT-5.5 on actually retrieving anything from it — a window you could fill but not use.
98.1% at 512k–1M closes that. The comparison point in Meta’s table is GPT-5.6 Sol at 73.8 in the same band, which is a 24-point margin at the hardest end of the range. If you have a workload that genuinely needs to reason over a very large corpus in a single pass — a full monorepo, a discovery set, a year of logs — this is the most credible option on the market this week, and it was not on the shortlist eight weeks ago.
That is a real reversal and worth stating plainly. It is also the reason the licensing question matters more now than it did in July. A mediocre closed model is easy to ignore. A leading one with a promise attached to it is a planning problem.
The weights that keep not arriving
The timeline is short enough to lay out in full.
- 9 July 2026 — Muse Spark 1.1 ships through the paid Meta Model API. Proprietary. This was the break from the open-weights identity Meta had spent years building.
- 5 August 2026 — Muse Spark 1.2 ships. Also closed, via the Meta Model API and OpenRouter.
- 10 August 2026 — Meta publishes Muse Glimmer, a 30-billion-parameter multimodal model under Apache 2.0, and Zuckerberg pledges that Muse Spark 1.2’s weights will open “in the coming weeks.”
- 2 September 2026 — Muse Spark 1.3 ships, proprietary. Zuckerberg says an open version is coming “soon.” Reporting describes 1.3’s weights as undecided.
- 3 September 2026 — Muse Spark 1.2’s weights have still not appeared in Meta’s Hugging Face organisations. Four weeks after “coming weeks,” the model they were promised for has already been superseded.
Note what the Glimmer release does inside that sequence. It is a genuine open-weight publication — Apache 2.0, small enough to run offline on a single 24GB GPU — and it arrived on the same day as the promise about the flagship. The open artifact was real; it just was not the model anyone was asking about. That is a familiar structure: ship the thing you are willing to give away, announce the thing you are not, and let the announcement carry the reputational weight.
We are not claiming Meta will never publish these weights. We are making a narrower and more useful point for anyone doing planning: across three consecutive flagship releases, the open-weight commitment has never once preceded the paid launch, and it has slipped every time. A roadmap promise with no date is not a supply you can architect against. If you need weights you can host — for cost control, for air-gapped deployment, for a compliance requirement, for insurance against the kind of mid-contract model access change OpenAI handed Cursor — build against something published. GLM 5.3 under MIT, Qwen under Apache 2.0, DeepSeek’s V4 line, Gemma 4, and Tencent’s HY4 all exist right now. That list has got dramatically stronger this quarter, which is the context in which Meta’s hesitation is most interesting.
What Meta thinks your data is worth
Here is the part almost nobody is reading, and it is on a public pricing page.
Muse Spark 1.3 has two endpoints:
| Standard | Contributor | |
|---|---|---|
| Input / 1M tokens | $1.25 | $0.10 |
| Output / 1M tokens | $4.25 | $0.20 |
| Cache read / 1M tokens | $0.15 | $0.002 |
| Your data | Private | ”Prompts and outputs may be used to improve Meta’s products” |
Take a mid-sized agentic workload: 100 million input tokens and 20 million output tokens in a month.
- Standard: (100 × $1.25) + (20 × $4.25) = $210
- Contributor: (100 × $0.10) + (20 × $0.20) = $14
A difference of $196 on 120 million tokens of traffic. That is $1.63 per million tokens, and it is the price Meta has publicly set on being allowed to train on your prompts and its own responses. Not an inferred figure, not an analyst estimate — arithmetic on the company’s own rate card. Cache reads make the point even more starkly: $0.15 against $0.002 is a 98.7% discount for the identical compute.
For comparison, the contributor rate sits below every serious commercial alternative. Gemini 3.8 Flash launched the day before at $0.75/$3.75 introductory — and doubles on 1 January 2027. Claude Sonnet 5 is $2/$10 standard. Meta’s contributor endpoint undercuts a cost-optimised small model by a wide margin while serving a model ranked sixth in the world. That price is not achievable on unit economics. It is achievable because you are paying the rest in kind.
This is a cleaner, more honest arrangement than most of the industry manages — the terms are stated, the price is posted, and the choice is yours to make. Compare it with the murk around what does and does not get retained on default consumer tiers, or with the carve-outs that exist only inside negotiated enterprise agreements. Meta has put a number on it. Our objection is not to the offer. It is that the offer is arriving in the space where the open weights were supposed to go.
Follow the substitution through. The original open-weights bargain was: Meta gives you the model, you give Meta an ecosystem, mindshare and free engineering. The 2026 bargain is: Meta rents you the model at a steep discount, and you give Meta your training data. Both are trades. The first one left you holding an asset. The second leaves you holding an invoice and a dependency — and the discount can be repriced at any renewal, in a way that a downloaded set of weights cannot.
Where “open” stopped buying anything
There is a common assumption that Meta’s open-weight posture is a regulatory play — that publishing weights buys relief under the EU AI Act. For a model in this class, it does not.
Article 53 does exempt genuinely free and open-source general-purpose models from parts of the technical documentation requirement. But the exemption collapses the moment a model is classified as carrying systemic risk: such a model owes every Article 53 obligation whatever licence it carries. A frontier model sitting sixth on a general intelligence index is not going to escape that classification.
So the licence choice here is a commercial decision wearing philosophical clothing, and it should be read that way. It is worth noting alongside the disclosure obligations under Article 50 that came into force last month — which apply to deployers like us, not just to labs — that the Act is increasingly indifferent to how a model is licensed and increasingly interested in what it can do. Openness is no longer a regulatory strategy. It has been reduced to what it always was for buyers: portability.
That is also the frame for the parallel move we described yesterday, when all three US frontier labs split their lineups into public and vetted tiers. Capability is becoming an entitlement attached to an account. Weights are the one form in which capability is not an entitlement — nobody can revoke a file you already have. That is precisely why the file keeps not arriving.
What to do
If you want the capability and the data is not sensitive: use the contributor endpoint and take the 93% saving. This is genuinely the cheapest frontier-class inference available today, and for public-data research, open-source work, synthetic data generation and prototyping there is no good argument against it.
If the data is sensitive: use the standard endpoint at $1.25/$4.25, which is competitive and keeps your prompts out of training. Make this a routing rule in code, not a habit — a single misdirected call cannot be recalled. And check whether the decision is even yours: if the data belongs to a customer under a processing agreement, your vendor’s training terms are their concern too.
If you are choosing a long-context model: 1.3 is now the one to beat, and the margin over GPT-5.6 Sol in the 512k–1M band is large enough to change real architectures. Test it on your own corpus.
If you are choosing a coding agent: benchmark it against your repository rather than against the table. The reported three-times increase in output tokens per response cuts against the reported 25% reduction in tokens per completed job, and which one dominates depends on whether your work is long agentic runs or short interactive edits. Our best AI coding tools guide covers where the incumbents currently earn their price, and Claude Code remains the reference point for terminal-native agent work.
If your plan depends on open weights: do not wait. Build against what is published and treat any Meta release as upside.
The bottom line
Muse Spark 1.3 is a good model and a better one than Meta has shipped before, particularly for long context, where it corrects the specific weakness we identified in its predecessor. That part is straightforwardly good news for buyers, and it puts real competitive pressure on the pricing of everything above it.
But the story of the release is what did not ship alongside it. Meta spent years building an identity on open weights, and has now gone three flagship releases in a row without publishing any — while promising each time that they are coming. In place of the weights there is now a posted price for your data: $1.63 per million tokens, take it or leave it, stated plainly enough that we can do the arithmetic in public.
The offer is fair on its own terms. The substitution is what to watch. “Open weights soon” has become the most reliably recurring item on Meta’s roadmap and the least reliably delivered, and a promise that has slipped three times should be priced as a promise, not as a plan.
Frequently asked questions
Is the contributor tier a good deal?
It depends entirely on what you are sending. The discount is real and large — $0.10/$0.20 per million tokens against $1.25/$4.25 on the standard endpoint, plus cache reads at $0.002 versus $0.15. For a workload running 100 million input and 20 million output tokens a month, that is roughly $14 against $210. Nothing else at this capability level is close to $0.10 input. But the consideration on the other side is not a privacy abstraction, it is a specific commitment: prompts and outputs may be used to improve Meta's products, and once traffic has gone through that endpoint you cannot retract it. The clean cases are the obvious ones. Public-data research, open-source code, synthetic data generation, throwaway prototyping, personal experimentation — send it, take the 93% saving. Customer records, internal source code, unreleased product material, anything covered by a client confidentiality obligation or a data processing agreement you signed with someone else — do not, at any price, and note that the decision is not yours to make unilaterally if a contract with a third party governs the data.
Should I switch from Claude or GPT to Muse Spark 1.3 for coding?
Run it on your own repository before you decide, because the benchmark picture and the cost picture point in different directions. On Meta's published numbers 1.3 leads Claude Opus 5 on DeepSWE v1.1 (75.4 to 74.0), Terminal-Bench 2.1 (88.8 to 86.7) and SWEAtlas CodeBase QnA (59.4 to 52.7), which is a serious result rather than a marketing one. But at least one independent write-up reports that 1.3 emits roughly three times the output tokens of 1.2 per response while needing about 25% fewer tokens to finish an entire coding job — those are compatible if the model thinks harder per step and takes fewer steps, but they resolve into very different invoices depending on whether your work is long-horizon agentic tasks or short interactive edits. Output tokens are the expensive side of every price sheet. Measure the cost of a completed unit of work, not the price per token.
When will the open weights actually be released?
Meta has not given a date, for either model. The Muse Spark 1.2 weights were pledged on 10 August 2026 alongside the Muse Glimmer release, described as arriving in the coming weeks, and as of 3 September they have not appeared in Meta's Hugging Face organisations. For 1.3 the company's position is softer still — reporting on the launch describes the weights as undecided, with Mark Zuckerberg promising an open version soon. The honest planning assumption is that a promised open-weight release is not a delivery date. If your architecture, your cost model or your compliance posture depends on weights you can host yourself, build against models whose weights exist today.
What open-weight models can I actually run right now instead?
Several, and the gap to the frontier is much narrower than it was a year ago. Muse Glimmer is Meta's own genuinely open release — 30 billion parameters under Apache 2.0, published on 10 August 2026 and small enough to run offline on a single 24GB GPU, though it is not in the same capability class as Muse Spark 1.3. Beyond Meta, the open-weight field has been the most active part of the market this quarter: Z.ai's GLM 5.3 family under MIT, Alibaba's Qwen line under Apache 2.0, DeepSeek's V4 variants, Google's Gemma 4, and Tencent's HY4. If you want the capability without hosting it, the standard Muse Spark endpoint at $1.25/$4.25 keeps your data private and is competitively priced against Gemini Flash — you just do not get to hold the weights.
Does releasing weights get Meta out of EU AI Act obligations?
Not for a model of this class, which is worth understanding because the assumption is common and wrong. Article 53 of the AI Act does exempt genuinely free and open-source general-purpose AI models from parts of the technical documentation burden. That exemption falls away entirely once a model is classified as carrying systemic risk — a model in that category owes every Article 53 obligation regardless of what licence it ships under. A frontier model sitting sixth on a general intelligence index is squarely in the territory where that classification applies. So whatever is driving the open-weights hesitation, regulatory relief is unlikely to be the prize on the other side. Meta also declined to sign the EU's voluntary code of practice in 2025, with policy chief Joel Kaplan saying Europe was heading down the wrong path on AI, so this is not a company that has been optimising for the Commission's good opinion.
Meta's benchmarks are Meta's benchmarks. How much should I trust them?
Treat the vendor table as a claim and the third-party placement as evidence. The independent data point is Artificial Analysis, which puts 1.3 sixth of 636 models on its Intelligence Index — that is a real result from an outside evaluator and it corroborates the general shape of Meta's argument, even though it does not confirm any individual score. Two specific cautions on the vendor table. First, the comparison against 1.2 pits 1.3's max reasoning mode against 1.2's xhigh, which is not a like-for-like configuration. Second, benchmark parameters can be gamed and frequently do not track how a model behaves on your own work. The long-context numbers are the ones we would weight most heavily here, because they are the ones that reverse a documented prior weakness rather than extend a strength.
What actually changed between Muse Spark 1.2 and 1.3?
Behaviour more than raw intelligence, and the behavioural changes are the ones that matter for agent work. Meta describes 1.3 as asking clarifying questions when a prompt is ambiguous, calling for help when it gets stuck, and confirming before it takes consequential actions. Those sound like soft product touches and are in fact the difference between an agent you can leave running and one you have to supervise. The efficiency claim attached to them is about 20% fewer tool calls and about 25% fewer tokens to complete a coding job. Meta also reports roughly a four-point improvement across benchmarks generally, and the model takes text, image and video input at a 1,048,576-token context window. It shipped simultaneously into Muse Code, Meta's terminal and CI coding agent, and the Meta Model API.
Sources
- Axios — Meta debuts Muse Spark 1.3 as personal agent work continues (2 September 2026)
- The Register — Zuck's Muse to Spark joy with open weights release 'soon' (2 September 2026)
- OpenRouter — Muse Spark 1.3 Contributor: API pricing and providers
- OpenRouter — Muse Spark 1.3: API pricing and providers
- Artificial Analysis — Muse Spark 1.3: intelligence, performance and price analysis
- OfficeChai — Meta releases Muse Spark 1.3, beats Opus 5 and GPT-5.6 Sol on some benchmarks
- CNBC — Meta launches Muse Glimmer open-weight AI model (10 August 2026)
- Constellation Research — Meta releases open weight Muse Glimmer model with open Muse Spark 1.2 on tap
- The Next Web — Meta's new AI model edges closer to OpenAI and Anthropic, its AI chief says
- eesel AI — Meta Muse Spark 1.3: benchmarks, pricing, and what changed
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.