Claude Opus 4.7 launched: what actually changed
Anthropic launched Claude Opus 4.7 on April 16, 2026 — two months after Opus 4.6. The headline numbers:
- SWE-bench Verified: 80.8% → 87.6% (+6.8 points)
- CursorBench: 58% → 70% (+12 points)
- Image resolution: 3x what Opus 4.6 could handle
- Pricing: unchanged at $5 per million input tokens / $25 per million output
That’s a real jump on coding benchmarks. Let me tell you what actually matters from two days of using it.
The SWE-bench jump is the real story
Everyone will quote the 87.6%. The number is defensible — Anthropic tests the standard 500-task SWE-bench Verified set, not a cherry-picked subset. Opus 4.7 is meaningfully better at solving real software engineering bugs than Opus 4.6 was.
That said, SWE-bench Verified is showing saturation. Scale AI’s newer SWE-bench Pro set — contamination-resistant, harder — puts top models around 45-50%, not 87%. The pragmatic read on Opus 4.7: on the kind of production bug-fixing work SWE-bench Verified models, it’s the best model shipping. On the hardest tasks (complex refactors, novel architecture), the gap between “best released model” and “best model in a lab” stays real.
What’s new beyond the benchmarks
High-resolution image support. First Claude model that can process images at over 3x the resolution of Opus 4.6. This matters if you’ve been frustrated by Claude blurring out diagrams, screenshots, or handwritten notes. It’s still not as good as GPT-5’s multimodal for spatial reasoning, but the gap is noticeably narrower.
Task budgets. Opus 4.7 introduces a new feature where you can give the model a rough token target for a full agentic loop — thinking, tool calls, tool results, final output. In practice: Claude will plan more conservatively when you say “solve this in 50K tokens” versus “take all day.” For agent workflows where cost is a real concern, this is the first Claude feature that lets you budget upfront instead of discovering the bill after.
Task-aware efficiency. Anthropic claims Opus 4.7 completes comparable agentic tasks with fewer tokens than 4.6, though they’re vague about the exact improvement. In my testing, the difference is noticeable on multi-turn conversations but not dramatic on single-shot tasks.
The Mythos revelation
The most interesting part of the announcement wasn’t what Anthropic released — it was what they said about what they haven’t released. Anthropic openly conceded Opus 4.7 trails an unreleased internal model called Claude Mythos Preview.
Mythos scores 93.9% on SWE-bench Verified, 94.6% on GPQA Diamond, and — unprecedentedly — reportedly found thousands of zero-day vulnerabilities across major operating systems and browsers during red-team evaluation. Anthropic is holding it back due to safety concerns.
Whatever you think of the framing, this is a first — a frontier lab publicly acknowledging its best model is too dangerous to ship. Combined with the CEO’s repeated warnings about capability growth, it suggests Anthropic is in an odd position: the commercial model it wants to sell is Opus 4.7; the model it knows exists is Mythos; the gap is getting uncomfortable to talk about.
What I’d recommend doing about it
If you pay for Claude Pro ($20): You get Opus 4.7. This is a real upgrade. The CursorBench jump especially matters for anyone using Claude for coding help.
If you pay for Max ($100 or $200): Same access with higher usage caps. Opus 4.7 is now the default for agentic work at Max.
If you use Claude Code: Opus 4.7 introduces a new xhigh effort tier between high and max. For most multi-file work it’s the new sweet spot — harder-thinking than high but not as expensive as max. Try it on your next real refactor.
If you’re on the API: Pricing is unchanged. Upgrade path from 4.6 to 4.7 should be a model-ID change. Watch for the tokenizer shift — Anthropic’s tokenizer change can raise effective costs by up to 35% even with unchanged sticker prices.
The bottom line
Opus 4.7 isn’t the generational leap Opus 4.0 was. It’s the steady, credible quarterly improvement Anthropic has been shipping since the Claude 3 era. On coding, it’s now genuinely ahead of GPT-5.4 and Gemini 3.1 Pro on the benchmarks most developers actually care about.
The more interesting question is what happens when Mythos ships — if it ever does. Anthropic is telling us the ceiling is much higher than what any of us can buy this month. That’s the real news.
Full Claude review, including Claude Code workflow recommendations, is at pick-right.com/tools/claude/.
Sources
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.