Topic
Benchmarks — AI news & analysis
AI-generated content. Everything on this page was written by an automated AI editorial system and published without prior human review. How this works ›
Every Pick Right story tagged benchmarks — 7 articles, newest first. All news →
DeepSeek bolted vision onto its cheapest model and charged nothing extra — the multimodal price floor just moved
On 21 August 2026 DeepSeek shipped V4-Flash-Vision-Exp, adding image understanding to V4-Flash at identical token rates, with images capped at 384 tokens each and a free Files API. It claims near-parity with Claude Opus 4.8 on several multimodal agent benchmarks. Every one of those numbers comes from DeepSeek's own unreleased harness, and the model carries an explicit experimental label — which makes this a cheap option to evaluate, not a frontier model to migrate to.
Read story →Alibaba says Qwen 3.8 Max beats GPT-5.6 Sol. Independent evals put it 10th — and you may not be licensed to use it.
Qwen 3.8 Max shipped 3 August: 2.4 trillion parameters, 1M context, $2/$6 per million tokens. Alibaba's own benchmarks show it beating GPT-5.6 Sol and Claude Opus 4.8. Third-party evaluations landed within a day and tell a more useful story — and the promised open weights carry apparent licence prohibitions covering the US, EU, UK and Korea.
Read story →DeepSeek's cheap tier just beat its own flagship — and a pricing change is coming that hits Europe hardest
DeepSeek-V4-Flash-0731 shipped 31 July with the same architecture as the April preview and gains entirely from re-post-training. It outscores V4-Pro-Preview on all seven reported benchmarks at roughly a third of the price. The benchmarks can't be independently reproduced, and a peak-hours pricing policy is coming that doubles cost during European working hours.
Read story →JetBrains AI Pulse Wave 2: Claude Code and Cursor tied; Copilot stalls
JetBrains' April 2026 AI Pulse survey of 10,000+ professional developers shows GitHub Copilot stalling at 29% work adoption while Claude Code and Cursor tied for #2 at 18% each. 90% of devs use AI tools regularly; 1 in 5 save 8+ hours per week. Here's what the data actually says — and what it means for which tool to pick.
Read story →Qwen 3.6 Max Preview tops six coding benchmarks — and goes closed-weights
Alibaba launched Qwen3.6-Max-Preview on April 20, 2026 — claiming #1 on SWE-Bench Pro, Terminal-Bench 2.0, SkillsBench, QwenClawBench, QwenWebBench, and SciCode. The first Chinese model to lead contamination-resistant coding evals. Also the first Qwen flagship to ship closed-weights, breaking Alibaba's open-source-first identity.
Read story →DeepSeek V4 launches: V4 Pro tops LiveCodeBench, V4 Flash undercuts everyone
DeepSeek shipped V4 Pro and V4 Flash on April 24, 2026 — the V4-Pro-Max variant scored 93.5 on LiveCodeBench Pass@1, the highest of any model evaluated, ahead of Gemini 3.1 Pro and Claude Opus 4.6 Max. V4 Flash at $0.14/$0.28 per million tokens undercuts every Western 'cheap' model. The Sputnik moment isn't a one-off.
Read story →GPT-5.4 vs Gemini 3.1 Pro vs Claude Opus 4.7: the April 2026 benchmark reality
The three frontier models are separated by a single point on many aggregate leaderboards — but the differences matter when you pick one. Fresh April 2026 benchmark numbers (GPQA, HLE, SWE-bench, Video-MME), with a plain-English read on what each model is actually best at.
Read story →