GPT-5.4 vs Gemini 3.1 Pro vs Claude Opus 4.7: the April 2026 benchmark reality
Three frontier models are within rounding error on top-line leaderboards this month. Zoom in and the differences get interesting — and actionable, if you’re choosing which $20-$200/month subscription to pay for.
The numbers
| Benchmark | GPT-5.4 | Gemini 3.1 Pro | Claude Opus 4.7 | Leader |
|---|---|---|---|---|
| BenchLM aggregate | 84 | 83 | — | GPT-5.4 (+1) |
| GPQA Diamond (science reasoning) | 92.0% | 94.1% | 92.5%* | Gemini 3.1 |
| Humanity’s Last Exam (April 19) | 41.6% | 44.7% | — | Gemini 3.1 |
| SWE-bench Verified (coding) | 78.2% | 78.8% | 87.6% | Opus 4.7 |
| CursorBench (coding agents) | — | — | 70% | Opus 4.7 (+12 vs 4.6) |
| SWE-Pro (realistic codebases) | 57.7 | 72.0 | — | Gemini 3.1 |
| HumanEval+ (code completion) | ~95% | 94% | — | GPT-5.4 (marginal) |
| Video-MME (multimodal) | 71.4% | 78.2% | — | Gemini 3.1 (+6.8) |
| API list price / 1M tokens | $2.50 / $20 | $2.00 /? | $5 / $25 | Gemini (cheapest) |
| Context window | 1M | 1M | 200K | Tie on 1M |
*Claude Opus 4.7 GPQA figure is Anthropic’s published number; leaderboards haven’t fully retested yet.
Sources: BenchLM, PricePerToken leaderboards, AI Magicx comparison.
What the numbers actually tell you
Gemini 3.1 Pro wins reasoning, science, multimodal, and long-context. The 94.1% GPQA Diamond is a new high-water mark. The 78.2% Video-MME score is 6.8 points clear of GPT-5.4, which is the biggest single-benchmark lead any frontier model has right now. Plus the 1M context at ~$2/M input tokens — roughly a third of Opus 4.7’s input price. If you’re doing research with citations, analyzing long documents, or processing video, Gemini 3.1 is genuinely the best model this week.
Claude Opus 4.7 wins coding — decisively. 87.6% on SWE-bench Verified is 8.8 points clear of the other two. CursorBench went 58% → 70%, which is the biggest one-quarter jump any coding model has posted. On the kind of real software engineering work Claude Code and Cursor use these models for, Opus 4.7 is the answer. Caveat: SWE-bench Verified is showing saturation. The new SWE-bench Pro set has top models at 45-50%, not 87%.
GPT-5.4 wins breadth. It doesn’t lead any single benchmark by much, but it’s rarely worse than second and it has the deepest ecosystem: ChatGPT, Code Interpreter, plugins, voice, image generation, Sora (until Saturday), and Codex. If you want one AI subscription for everything, GPT-5.4 via ChatGPT Plus is still the right pick.
The benchmark caveats everyone should know
Verified sets are contaminated. SWE-bench Verified was groundbreaking in 2024. By 2026, all three frontier labs have almost certainly seen variants of the test tasks during training. The SWE-bench Pro leaderboard shows top models scoring 45-50% — the current ceiling on hard coding. Expect Pro to become the default benchmark cited in serious reviews over the next six months.
HLE is not saturated yet. Humanity’s Last Exam is still hard — even the best models score below 50%. The 41-45% range represents genuine remaining hard reasoning tasks. Gemini’s lead here is real signal, not test contamination.
“Aggregate leaderboard” scores are marketing. BenchLM’s 84 vs 83 for GPT-5.4 vs Gemini is statistical noise. Real decisions should come from specific benchmarks matched to your specific use case.
Benchmark ≠ product. Opus 4.7 winning SWE-bench doesn’t mean Claude is the best product for every developer. Opus has 200K context vs Gemini/GPT’s 1M — for huge codebases, Gemini 3.1 can reason over everything Opus 4.7 physically can’t see.
What I’d actually recommend in April 2026
Default for most users: ChatGPT Plus ($20/mo). GPT-5.4 is rarely the wrong answer, and the ecosystem breadth (voice, images, GPTs, Code Interpreter) is unmatched.
If you write or code professionally: Claude Pro ($20/mo). Opus 4.7 is meaningfully ahead on writing quality and coding, and Claude Code is the best agentic coding workflow.
If you do research or need citations: Google AI Pro ($19.99/mo). Gemini 3.1 Pro’s Deep Search is the most-trusted research output, plus 2TB Drive makes the bundle effectively free if you already pay for storage.
If you’re a heavy API user: Gemini 3.1 Pro at $2/M input is dramatically cheaper than Claude Opus 4.7 at $5/M. For high-volume automation where model quality is similar, Gemini wins on economics.
The working stack for people who can afford it: ChatGPT Plus + Claude Pro + Google AI Pro = $60/month. Each does one thing better than the others. The composite “best AI assistant” is all three, switched between based on task.
Full per-tool breakdowns:
Or jump to the ranked category pages:
Sources
Related tool reviews
Questions or corrections? Email Pick Right. Want the full list? See all news.