FrontierCode
Cognition's benchmark for code mergeability and production quality, not just correctness. Tasks are drawn from 36 real open-source repositories and authored by their maintainers (40+ hours each), with concise, humanlike prompts (~1/3 the length of SWE-bench Pro). Solutions are graded against a maintainer-style rubric spanning behavioral correctness, regression safety, mechanical cleanliness, test correctness, scope, and code quality; the reported score is a weighted aggregate of the rubric items, and any solution that fails a 'blocker' criterion scores 0. Revision 1.1 (2026-07-07) publishes two nested subsets — Main (100 tasks) and Extended (150) — having deprecated the original Diamond (50 hardest) subset, and zeroes runs flagged for consulting solution-bearing sources such as the original pull request. Each model is run 5× at every available reasoning effort and its best effort is reported. Tasks are kept private to avoid contamination.
What is FrontierCode?
Cognition's benchmark for code mergeability and production quality, not just correctness. Tasks are drawn from 36 real open-source repositories and authored by their maintainers (40+ hours each), with concise, humanlike prompts (~1/3 the length of SWE-bench Pro). Solutions are graded against a maintainer-style rubric spanning behavioral correctness, regression safety, mechanical cleanliness, test correctness, scope, and code quality; the reported score is a weighted aggregate of the rubric items, and any solution that fails a 'blocker' criterion scores 0. Revision 1.1 (2026-07-07) publishes two nested subsets — Main (100 tasks) and Extended (150) — having deprecated the original Diamond (50 hardest) subset, and zeroes runs flagged for consulting solution-bearing sources such as the original pull request. Each model is run 5× at every available reasoning effort and its best effort is reported. Tasks are kept private to avoid contamination. evals.report tracks reported FrontierCode scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.
Top reported FrontierCode score: Claude Fable 5 — 53.5% (weighted score (Main)).
| Model | Lab | Score↓ | Source model | Status | Date | |
|---|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | 53.5% | Fable 5 | Official | Jun 9, 2026 | Details |
| Claude Opus 5 | Anthropic | 53.4% | Opus 5 | Official | Jul 24, 2026 | Details |
| GPT-6 Astra | OpenAI | 53.3% | GPT-6 Astra | Official | Sep 3, 2026 | Details |
| Claude Fable 5.1 | Anthropic | 50.9% | Fable 5.1 | Official | Sep 1, 2026 | Details |
| Grok 4.6 | xAI | 48.0% | Grok 4.6 | Official | Aug 12, 2026 | Details |
| GPT-5.6 Sol | OpenAI | 47.5% | GPT-5.6 Sol | Official | Jul 9, 2026 | Details |
| Claude Opus 4.8 | Anthropic | 46.5% | Opus 4.8 | Official | May 28, 2026 | Details |
| Kimi K3Open | Moonshot AI | 44.2% | Kimi K3 | Official | Jul 16, 2026 | Details |
| Gemini 3.7 Flash | Google DeepMind | 43.6% | Gemini 3.7 Flash | Verified | Aug 13, 2026 | Details |
| GPT-5.5 | OpenAI | 43.0% | GPT-5.5 | Official | Apr 23, 2026 | Details |
| Claude Sonnet 5 | Anthropic | 42.7% | Sonnet 5 | Official | Jun 30, 2026 | Details |
| Grok 4.5 | xAI | 42.4% | Grok 4.5 | Official | Jul 8, 2026 | Details |
| SWE-1.7 | Cognition | 42.0% | SWE-1.7 | Official | Jul 8, 2026 | Details |
| GPT-5.6 Terra | OpenAI | 41.3% | GPT-5.6 Terra | Official | Jul 9, 2026 | Details |
| Gemini 3.8 Flash | Google DeepMind | 41.2% | Gemini 3.8 Flash | Official | Sep 2, 2026 | Details |
| GLM-5.3Open | Z.ai | 40.1% | GLM 5.3 | Official | Aug 14, 2026 | Details |
| GPT-5.6 Luna | OpenAI | 39.8% | GPT-5.6 Luna | Official | Jul 9, 2026 | Details |
| Claude Opus 4.7 | Anthropic | 38.5% | Opus 4.7 | Official | Apr 16, 2026 | Details |
| Gemini 3.6 Flash | Google DeepMind | 34.4% | Gemini 3.6 Flash | Verified | Jul 21, 2026 | Details |
| GLM-5.3-FlashOpen | Z.ai | 31.8% | GLM 5.3 Flash | Official | Aug 26, 2026 | Details |
| Kimi K2.7Open | Moonshot AI | 30.1% | Kimi K2.7 | Official | — | Details |
| DeepSeek V4 Pro 0813Open | DeepSeek | 28.5% | DeepSeek V4 Pro 0813 | Official | Aug 13, 2026 | Details |
| GPT-5.4-mini | OpenAI | 27.0% | GPT-5.4-mini | Official | Mar 17, 2026 | Details |
| Claude Opus 4.6 | Anthropic | 26.6% | Opus 4.6 | Official | Feb 5, 2026 | Details |
| Composer 2.5 | Cursor | 25.6% | Composer 2.5 | Official | — | Details |
| GLM-5.2Open | Z.ai | 24.5% | GLM 5.2 | Official | Jun 16, 2026 | Details |
| Claude Sonnet 4.6 | Anthropic | 24.3% | Sonnet 4.6 | Official | Feb 17, 2026 | Details |
| DeepSeek V4 Flash 0731Open | DeepSeek | 18.8% | DeepSeek V4 Flash 0731 | Official | Jul 31, 2026 | Details |
| DeepSeek V4 ProOpen | DeepSeek | 17.6% | DeepSeek V4 Pro | Official | Apr 24, 2026 | Details |
| MiniMax M3Open | MiniMax | 14.7% | MiniMax M3 | Official | Jun 1, 2026 | Details |
| InklingOpen | Thinking Machines Lab | 14.0% | Inkling | Official | Jul 15, 2026 | Details |
| NVIDIA Nemotron 3 UltraOpen | NVIDIA | 13.6% | Nemotron 3 Ultra | Official | Jun 4, 2026 | Details |
| Qwen3.7 Plus | Alibaba / Qwen | 10.2% | Qwen 3.7 Plus | Official | — | Details |
| SWE-1.6 | Cognition | 9.4% | SWE-1.6 | Official | Apr 7, 2026 | Details |
| Mistral Medium 3.5 | Mistral AI | 8.0% | Mistral 3.5 Medium | Official | Apr 28, 2026 | Details |
Each row reports the model’s weighted score (Main) on FrontierCode. Click a row for the full run context.