Some links on this page are affiliate links. We may earn a commission at no extra cost to you.
Updated: Jul 21, 2026
·
metamodelsagents

Meta's Muse Spark 1.1 quietly ends the open-only era — a cheap agentic model that leads on tool use and stumbles on long context

TL;DR: Meta Superintelligence Labs shipped Muse Spark 1.1 on July 9 through the paid Meta Model API and inside Meta AI — a quiet but real break for the company that made its name on open weights. Several outlets call it Meta’s first paid agent model. It’s cheap ($1.25 / $4.25 per million tokens) and genuinely strong where it counts for agents: it leads tool-use benchmarks (JobBench 54.7, MCP Atlas 88.1) and tops Humanity’s Last Exam with and without tools. The honest weaknesses: it trails on pure coding (SWE-Bench Pro 61.5 vs Claude Opus 4.8’s 69.2; DeepSWE 53.3), and despite a 1M-token context window, Meta’s own table shows MRCR long-context retrieval at 54.1 against GPT-5.5’s 74.0. What this means for you: a strong, cheap pick for tool-calling agents — not the pick for production coding or long-document recall.

What shipped

On July 9, 2026, Meta Superintelligence Labs released Muse Spark 1.1, a multimodal reasoning model built for agentic tasks. The distribution is the story as much as the model:

For a company whose AI identity was built on giving Llama’s weights away, shipping a frontier-tier agentic model behind a paid API is a genuine pivot. The Agent Report and others frame it as Meta’s first paid agent model. That framing is fair, with one qualification worth keeping: this is the end of open-only, not necessarily the end of open weights at Meta.

The benchmark picture — read it carefully

Muse Spark 1.1’s results are unusually lopsided, and that’s genuinely useful information rather than a flaw in the reporting.

Where it leads:

Where it trails:

That last one deserves emphasis because it’s the number a spec sheet will never show you. A 1M-token context window and reliable retrieval across 1M tokens are different claims. The window is how much you can put in; the retrieval score is how much the model can actually find and use once it’s in there. A 20-point gap to GPT-5.5 on that axis means the headline context number oversells what you’ll get in practice for long-document work.

Why this matters

1. Meta’s open-only era is over, and that reshapes the open-weight map. Meta was the credibility anchor of Western open weights. With its best agentic model behind a paid API, the open-weight frontier is now led substantially by Chinese labs — Kimi K3, GLM-5.2, DeepSeek, Qwen. That’s a meaningful shift in who sets the terms of open AI, and it’s worth noticing regardless of which model you use.

2. Specialisation is beating generality, and the benchmarks are finally showing it honestly. Muse Spark 1.1 isn’t trying to be the best at everything, and its scorecard says so plainly: excellent at tool use, mid-tier at coding, weak at long-context recall. That’s more useful to a buyer than a model that claims to lead everywhere. The 2026 pattern is clear — Grok 4.5 wins on cost-per-task, Kimi K3 wins frontend code, Muse Spark wins tool use — while the generalist crown stays with Claude and GPT-5.6.

3. Cheap tool-calling is a real unlock. Agentic workloads burn tokens: plan, call, read result, re-plan, call again. At $1.25/$4.25 with leading tool-use scores, Muse Spark 1.1 targets exactly the workload where token cost compounds fastest. If you’ve priced out an agent on frontier models and flinched, this is worth an evaluation — it’s the clearest use case the model has.

4. It sharpens the “don’t trust the context number” lesson. We’ve made this argument about benchmarks generally — METR found GPT-5.6 Sol games its own evaluations — and Muse Spark supplies the context-window version. Vendors advertise window size because it’s a single impressive number. Retrieval quality across that window is the thing that determines whether your long-document pipeline works. Ask for the second number; if it isn’t published, assume it’s unflattering.

5. Meta is competing on price and distribution, not on the crown. Undercutting on price while shipping inside Meta AI’s enormous consumer surface is a distribution play, not a benchmark play. That’s a rational strategy for a company that’s behind on the frontier — and it adds another notch of pricing pressure on GPT-5.6’s Terra/Luna tiers and Claude’s mid-range.

Where it fits against the alternatives

The lopsided scorecard makes this an unusually easy model to place. Sorted by what you’re actually trying to do:

The broader lesson is that “which model is best” has become close to meaningless as a question. In 2026 the frontier labs hold the generalist crown while a widening field of specialists takes individual categories — often at a quarter of the price. The buyers getting the most value are the ones routing different workloads to different models rather than standardising on one.

What this means for you

The honest caveats

The grounded summary: Muse Spark 1.1 is a cheap, genuinely strong tool-use model with two clearly documented weaknesses — and a strategic marker that Meta’s open-only chapter has closed. Use it where it’s strong, and don’t let the 1M context number do work the retrieval score doesn’t support.

Frequently asked questions

What is Muse Spark 1.1?

Muse Spark 1.1 is a multimodal reasoning model built for agentic tasks, released by Meta Superintelligence Labs on July 9, 2026. It's available through the paid Meta Model API and inside Meta AI, with a 1 million-token context window and active context management. API pricing is $1.25 per million input tokens and $4.25 per million output tokens.

Why is a paid Meta model significant?

Because Meta built its AI reputation on releasing open weights — Llama was the standard-bearer for the open-weight movement. Shipping a frontier-tier agentic model behind a paid API instead marks a real strategic turn, described by several outlets as Meta's first paid agent model. It doesn't mean Meta has abandoned open weights, but it does mean the open-only identity is over.

What is Muse Spark 1.1 actually good at?

Agentic tool use is its clear strength. It leads the tool-use and tool-augmented reasoning benchmarks — JobBench at 54.7 and MCP Atlas at 88.1 — and tops Humanity's Last Exam both with and without tools. If your workload is an agent that calls tools, plans multi-step work, and reasons over results, this is genuinely competitive at a fraction of frontier pricing.

Where does it fall short?

Two places. Pure coding: SWE-Bench Pro 61.5 versus Claude Opus 4.8's 69.2, and DeepSWE 53.3 — respectable but not leading. And long-context retrieval: despite the 1M-token window, Meta's own table shows MRCR Long Context at 54.1 against GPT-5.5's 74.0. A big context window and reliable recall across that window are different things, and the gap here is large.

Should I use it instead of Claude or GPT-5.6?

Use it for tool-heavy agentic workloads where cost matters — that's where it's strongest and cheapest. For production coding, Claude Opus 4.8 and GPT-5.6 Sol still lead on the coding benchmarks. And if your use depends on reliably retrieving detail from very long documents, benchmark it yourself first; the 1M context number oversells the retrieval reality.

Sources

Related tool reviews

Questions or corrections? Email Pick Right. Want the full list? See all news.