
Four AIs hit 42/42 on IMO 2026.…
Claude Fable 5, GPT-5.6 Sol, Kimi K3, and AxiomProver all reported perfect IMO 2026 scores. The interesting part is cost, grading tier, and what happens when benchmarks stop separating models.

Claude Fable 5, GPT-5.6 Sol, Kimi K3, and AxiomProver all reported perfect IMO 2026 scores. The interesting part is cost, grading tier, and what happens when benchmarks stop separating models.

Tensorlake ran 30 hard agentic tasks with DeepSeek V4 Flash wired through four harnesses. Pi won on pass rate and cost per success. Claude Code was fastest but burned 741k tokens per task.

Cognition's FrontierCode benchmark grades mergeability, not just test passes. On the hardest Diamond tier, Claude Opus 4.8 leads at 13.4%. Here is how I read that number for production agent routing.

Datacurve's DeepSWE benchmark uses 113 original long-horizon tasks and hand-written verifiers so GPT-5.5 leads by 16 points where older SWE tests looked tied. Here is why that matters for model pickers.