BenchmarksReasoning
AIME 2026
Accuracy of LLMs on the 30 problems of the 2026 American Invitational Mathematics Examination (AIME I and II), a contamination-free competition-math benchmark requiring integer answers (0-999), evaluated live by MathArena.
ReasoningaccuracyHigher is better
| Model | Lab | Score↓ | Source model | Status | Date | |
|---|---|---|---|---|---|---|
| Claude Opus 4.8 | Anthropic | 100.00% | — | Official | May 28, 2026 | Details |
| GPT-5.4 | OpenAI | 99.17% | — | Official | Mar 5, 2026 | Details |
| GPT-5.2 | OpenAI | 98.33% | — | Official | Dec 11, 2025 | Details |
| Gemini 3.1 Pro Preview | Google DeepMind | 98.33% | — | Official | Feb 19, 2026 | Details |
| GPT-5.5 | OpenAI | 97.50% | — | Official | Apr 23, 2026 | Details |
| Claude Opus 4.6 | Anthropic | 96.67% | — | Official | Feb 5, 2026 | Details |
| DeepSeek V4 Flash | DeepSeek | 95.83% | — | Official | Apr 24, 2026 | Details |
| Gemini 3 Flash | Google DeepMind | 95.83% | — | Official | Dec 17, 2025 | Details |
| Kimi K2.5 | Moonshot AI | 95.83% | — | Official | Jan 27, 2026 | Details |
| GLM-5 | Z.ai | 95.83% | — | Official | Feb 11, 2026 | Details |
| DeepSeek V4 Pro | DeepSeek | 95.83% | — | Official | Apr 24, 2026 | Details |
| GLM-5.1 | Z.ai | 95.83% | — | Official | Apr 7, 2026 | Details |
| Kimi K2.6 | Moonshot AI | 95.83% | — | Official | Apr 20, 2026 | Details |
| Claude Opus 4.7 | Anthropic | 95.83% | — | Official | Apr 16, 2026 | Details |
| Gemini 3.5 Flash | Google DeepMind | 95.00% | — | Official | May 19, 2026 | Details |
| Grok 4.1 fast reasoning | xAI | 94.17% | — | Official | Nov 19, 2025 | Details |
| DeepSeek V3.2 | DeepSeek | 94.17% | — | Official | Dec 1, 2025 | Details |
| Qwen3.5-397B-A17B | Alibaba / Qwen | 93.33% | — | Official | Feb 16, 2026 | Details |
| Gemini 3 Pro | Google DeepMind | 91.67% | — | Official | Nov 18, 2025 | Details |
| NVIDIA Nemotron 3 Super 120B-A12B | NVIDIA | 90.00% | — | Official | Mar 10, 2026 | Details |
Each row reports the model’s accuracy on AIME 2026. Click a row for the full run context.