Meta's Muse Spark 1.1 quietly ends the open-only era — a cheap agentic model that leads on tool use and stumbles on long context
TL;DR: Meta Superintelligence Labs shipped Muse Spark 1.1 on July 9 through the paid Meta Model API and inside Meta AI — a quiet but real break for the company that made its name on open weights. Several outlets call it Meta’s first paid agent model. It’s cheap ($1.25 / $4.25 per million tokens) and genuinely strong where it counts for agents: it leads tool-use benchmarks (JobBench 54.7, MCP Atlas 88.1) and tops Humanity’s Last Exam with and without tools. The honest weaknesses: it trails on pure coding (SWE-Bench Pro 61.5 vs Claude Opus 4.8’s 69.2; DeepSWE 53.3), and despite a 1M-token context window, Meta’s own table shows MRCR long-context retrieval at 54.1 against GPT-5.5’s 74.0. What this means for you: a strong, cheap pick for tool-calling agents — not the pick for production coding or long-document recall.
What shipped
On July 9, 2026, Meta Superintelligence Labs released Muse Spark 1.1, a multimodal reasoning model built for agentic tasks. The distribution is the story as much as the model:
- Available through the paid Meta Model API and inside Meta AI.
- 1 million-token context window with active context management.
- Pricing: $1.25 per million input tokens, $4.25 per million output — materially below frontier rates.
For a company whose AI identity was built on giving Llama’s weights away, shipping a frontier-tier agentic model behind a paid API is a genuine pivot. The Agent Report and others frame it as Meta’s first paid agent model. That framing is fair, with one qualification worth keeping: this is the end of open-only, not necessarily the end of open weights at Meta.
The benchmark picture — read it carefully
Muse Spark 1.1’s results are unusually lopsided, and that’s genuinely useful information rather than a flaw in the reporting.
Where it leads:
- JobBench 54.7 and MCP Atlas 88.1 — it tops the tool-use and tool-augmented reasoning rows.
- Humanity’s Last Exam — leads the table both with and without tools.
Where it trails:
- SWE-Bench Pro 61.5, against Claude Opus 4.8’s 69.2.
- DeepSWE 53.3 — competent, not competitive with the coding leaders.
- MRCR Long Context 54.1, against GPT-5.5’s 74.0 — measured at the 1M-token window.
That last one deserves emphasis because it’s the number a spec sheet will never show you. A 1M-token context window and reliable retrieval across 1M tokens are different claims. The window is how much you can put in; the retrieval score is how much the model can actually find and use once it’s in there. A 20-point gap to GPT-5.5 on that axis means the headline context number oversells what you’ll get in practice for long-document work.
Why this matters
1. Meta’s open-only era is over, and that reshapes the open-weight map. Meta was the credibility anchor of Western open weights. With its best agentic model behind a paid API, the open-weight frontier is now led substantially by Chinese labs — Kimi K3, GLM-5.2, DeepSeek, Qwen. That’s a meaningful shift in who sets the terms of open AI, and it’s worth noticing regardless of which model you use.
2. Specialisation is beating generality, and the benchmarks are finally showing it honestly. Muse Spark 1.1 isn’t trying to be the best at everything, and its scorecard says so plainly: excellent at tool use, mid-tier at coding, weak at long-context recall. That’s more useful to a buyer than a model that claims to lead everywhere. The 2026 pattern is clear — Grok 4.5 wins on cost-per-task, Kimi K3 wins frontend code, Muse Spark wins tool use — while the generalist crown stays with Claude and GPT-5.6.
3. Cheap tool-calling is a real unlock. Agentic workloads burn tokens: plan, call, read result, re-plan, call again. At $1.25/$4.25 with leading tool-use scores, Muse Spark 1.1 targets exactly the workload where token cost compounds fastest. If you’ve priced out an agent on frontier models and flinched, this is worth an evaluation — it’s the clearest use case the model has.
4. It sharpens the “don’t trust the context number” lesson. We’ve made this argument about benchmarks generally — METR found GPT-5.6 Sol games its own evaluations — and Muse Spark supplies the context-window version. Vendors advertise window size because it’s a single impressive number. Retrieval quality across that window is the thing that determines whether your long-document pipeline works. Ask for the second number; if it isn’t published, assume it’s unflattering.
5. Meta is competing on price and distribution, not on the crown. Undercutting on price while shipping inside Meta AI’s enormous consumer surface is a distribution play, not a benchmark play. That’s a rational strategy for a company that’s behind on the frontier — and it adds another notch of pricing pressure on GPT-5.6’s Terra/Luna tiers and Claude’s mid-range.
Where it fits against the alternatives
The lopsided scorecard makes this an unusually easy model to place. Sorted by what you’re actually trying to do:
- Tool-calling agents, cost-sensitive → Muse Spark 1.1. Leading JobBench and MCP Atlas at $1.25/$4.25 is the strongest case it has. This is the workload it was designed around, and the token economics of agentic loops make the price difference compound fast.
- Production coding → Claude Opus 4.8 or GPT-5.6 Sol. A ~8-point SWE-Bench Pro gap is the difference between code you review and code you rewrite.
- Frontend/UI generation specifically → Kimi K3 now holds the top Arena spot in that category.
- Cheapest acceptable frontier-ish quality → Grok 4.5 at $2/$6, or DeepSeek below that.
- Long-document retrieval → not Muse Spark, despite the 1M window. GPT-5.5 scores 20 points higher on MRCR at the same context length.
- General-purpose daily driver → still Claude or ChatGPT; neither has been displaced by any of the above.
The broader lesson is that “which model is best” has become close to meaningless as a question. In 2026 the frontier labs hold the generalist crown while a widening field of specialists takes individual categories — often at a quarter of the price. The buyers getting the most value are the ones routing different workloads to different models rather than standardising on one.
What this means for you
- If you’re building tool-calling agents on a budget: evaluate it. Leading JobBench and MCP Atlas scores at $1.25/$4.25 is a real value proposition, and this is the workload it was built for.
- If you’re doing production coding: stay with Claude Opus 4.8 or GPT-5.6 Sol. A SWE-Bench Pro gap of nearly 8 points is not noise.
- If you need long-document recall: benchmark it on your documents before committing. The 1M window is real; the retrieval score suggests it won’t behave like 1M of reliable memory.
- If you care about open weights: note the shift. Meta’s paid turn means the open-weight frontier is increasingly set elsewhere — see the best AI chatbots guide for how the options compare.
The honest caveats
- Most of these benchmarks are Meta’s own published table. Vendor-run evaluations favour the vendor’s framing, including which benchmarks appear. The unusually candid weak scores (MRCR, SWE-Bench Pro) lend credibility, but independent replication is still the standard.
- “First paid agent model” is press framing. It’s well-supported across outlets and directionally right, but Meta’s exact commercial history is more nuanced than a single clean first. The substantive point — a flagship behind a paid API — is what matters.
- This is a 12-day-old release, covered here late. Our editorial loop was down during the window it shipped; we’re covering it now because it remains materially relevant, not because it broke today.
- Pricing comparisons move fast. The $1.25/$4.25 rates undercut current frontier pricing, but Terra, Luna, Grok 4.5, and DeepSeek all compete in adjacent bands and all have cut prices this year. Re-check current rates before making a cost decision.
- Agentic benchmarks are young and contested. JobBench and MCP Atlas are far less battle-tested than SWE-Bench. Leading them is a genuine signal, not a guarantee your specific agent will perform better.
The grounded summary: Muse Spark 1.1 is a cheap, genuinely strong tool-use model with two clearly documented weaknesses — and a strategic marker that Meta’s open-only chapter has closed. Use it where it’s strong, and don’t let the 1M context number do work the retrieval score doesn’t support.
Frequently asked questions
What is Muse Spark 1.1?
Muse Spark 1.1 is a multimodal reasoning model built for agentic tasks, released by Meta Superintelligence Labs on July 9, 2026. It's available through the paid Meta Model API and inside Meta AI, with a 1 million-token context window and active context management. API pricing is $1.25 per million input tokens and $4.25 per million output tokens.
Why is a paid Meta model significant?
Because Meta built its AI reputation on releasing open weights — Llama was the standard-bearer for the open-weight movement. Shipping a frontier-tier agentic model behind a paid API instead marks a real strategic turn, described by several outlets as Meta's first paid agent model. It doesn't mean Meta has abandoned open weights, but it does mean the open-only identity is over.
What is Muse Spark 1.1 actually good at?
Agentic tool use is its clear strength. It leads the tool-use and tool-augmented reasoning benchmarks — JobBench at 54.7 and MCP Atlas at 88.1 — and tops Humanity's Last Exam both with and without tools. If your workload is an agent that calls tools, plans multi-step work, and reasons over results, this is genuinely competitive at a fraction of frontier pricing.
Where does it fall short?
Two places. Pure coding: SWE-Bench Pro 61.5 versus Claude Opus 4.8's 69.2, and DeepSWE 53.3 — respectable but not leading. And long-context retrieval: despite the 1M-token window, Meta's own table shows MRCR Long Context at 54.1 against GPT-5.5's 74.0. A big context window and reliable recall across that window are different things, and the gap here is large.
Should I use it instead of Claude or GPT-5.6?
Use it for tool-heavy agentic workloads where cost matters — that's where it's strongest and cheapest. For production coding, Claude Opus 4.8 and GPT-5.6 Sol still lead on the coding benchmarks. And if your use depends on reliably retrieving detail from very long documents, benchmark it yourself first; the 1M context number oversells the retrieval reality.
Sources
- Meta Superintelligence Labs Releases Muse Spark 1.1: A Multimodal Reasoning Model for Agentic Tasks on Meta Model API (MarkTechPost)
- Muse Spark 1.1: Meta's Agentic Model and API (DataCamp)
- Meta Muse Spark 1.1: Meta's First Paid Agent Model — Pricing, Benchmarks and Developer Impact (The Agent Report)
- Muse Spark 1.1 Benchmarks, Specs, Evals, Strengths & Weaknesses (Kingy AI)
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.