evals.report
BenchmarksLabsCompareRun guidesIn the wild

FrontierCode

Cognition's benchmark for code mergeability and production quality, not just correctness. Tasks are drawn from 36 real open-source repositories and authored by their maintainers (40+ hours each), with concise, humanlike prompts (~1/3 the length of SWE-bench Pro). Solutions are graded against a maintainer-style rubric spanning behavioral correctness, regression safety, mechanical cleanliness, test correctness, scope, and code quality; the reported score is a weighted aggregate of the rubric items, and any solution that fails a 'blocker' criterion scores 0. Revision 1.1 (2026-07-07) publishes two nested subsets — Main (100 tasks) and Extended (150) — having deprecated the original Diamond (50 hardest) subset, and zeroes runs flagged for consulting solution-bearing sources such as the original pull request. Each model is run 5× at every available reasoning effort and its best effort is reported. Tasks are kept private to avoid contamination.

Codingweighted score (Main)Higher is better

What is FrontierCode?

Cognition's benchmark for code mergeability and production quality, not just correctness. Tasks are drawn from 36 real open-source repositories and authored by their maintainers (40+ hours each), with concise, humanlike prompts (~1/3 the length of SWE-bench Pro). Solutions are graded against a maintainer-style rubric spanning behavioral correctness, regression safety, mechanical cleanliness, test correctness, scope, and code quality; the reported score is a weighted aggregate of the rubric items, and any solution that fails a 'blocker' criterion scores 0. Revision 1.1 (2026-07-07) publishes two nested subsets — Main (100 tasks) and Extended (150) — having deprecated the original Diamond (50 hardest) subset, and zeroes runs flagged for consulting solution-bearing sources such as the original pull request. Each model is run 5× at every available reasoning effort and its best effort is reported. Tasks are kept private to avoid contamination. evals.report tracks reported FrontierCode scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.

Top reported FrontierCode score: Claude Fable 5 53.5% (weighted score (Main)).

ModelLabScoreSource modelStatusDate
Claude Fable 5Anthropic53.5%Fable 5OfficialJun 9, 2026Details
Claude Opus 5Anthropic53.4%Opus 5OfficialJul 24, 2026Details
GPT-6 AstraOpenAI53.3%GPT-6 AstraOfficialSep 3, 2026Details
Claude Fable 5.1Anthropic50.9%Fable 5.1OfficialSep 1, 2026Details
Grok 4.6xAI48.0%Grok 4.6OfficialAug 12, 2026Details
GPT-5.6 SolOpenAI47.5%GPT-5.6 SolOfficialJul 9, 2026Details
Claude Opus 4.8Anthropic46.5%Opus 4.8OfficialMay 28, 2026Details
Kimi K3OpenMoonshot AI44.2%Kimi K3OfficialJul 16, 2026Details
Gemini 3.7 FlashGoogle DeepMind43.6%Gemini 3.7 FlashVerifiedAug 13, 2026Details
GPT-5.5OpenAI43.0%GPT-5.5OfficialApr 23, 2026Details
Claude Sonnet 5Anthropic42.7%Sonnet 5OfficialJun 30, 2026Details
Grok 4.5xAI42.4%Grok 4.5OfficialJul 8, 2026Details
SWE-1.7Cognition42.0%SWE-1.7OfficialJul 8, 2026Details
GPT-5.6 TerraOpenAI41.3%GPT-5.6 TerraOfficialJul 9, 2026Details
Gemini 3.8 FlashGoogle DeepMind41.2%Gemini 3.8 FlashOfficialSep 2, 2026Details
GLM-5.3OpenZ.ai40.1%GLM 5.3OfficialAug 14, 2026Details
GPT-5.6 LunaOpenAI39.8%GPT-5.6 LunaOfficialJul 9, 2026Details
Claude Opus 4.7Anthropic38.5%Opus 4.7OfficialApr 16, 2026Details
Gemini 3.6 FlashGoogle DeepMind34.4%Gemini 3.6 FlashVerifiedJul 21, 2026Details
GLM-5.3-FlashOpenZ.ai31.8%GLM 5.3 FlashOfficialAug 26, 2026Details
Kimi K2.7OpenMoonshot AI30.1%Kimi K2.7OfficialDetails
DeepSeek V4 Pro 0813OpenDeepSeek28.5%DeepSeek V4 Pro 0813OfficialAug 13, 2026Details
GPT-5.4-miniOpenAI27.0%GPT-5.4-miniOfficialMar 17, 2026Details
Claude Opus 4.6Anthropic26.6%Opus 4.6OfficialFeb 5, 2026Details
Composer 2.5Cursor25.6%Composer 2.5OfficialDetails
GLM-5.2OpenZ.ai24.5%GLM 5.2OfficialJun 16, 2026Details
Claude Sonnet 4.6Anthropic24.3%Sonnet 4.6OfficialFeb 17, 2026Details
DeepSeek V4 Flash 0731OpenDeepSeek18.8%DeepSeek V4 Flash 0731OfficialJul 31, 2026Details
DeepSeek V4 ProOpenDeepSeek17.6%DeepSeek V4 ProOfficialApr 24, 2026Details
MiniMax M3OpenMiniMax14.7%MiniMax M3OfficialJun 1, 2026Details
InklingOpenThinking Machines Lab14.0%InklingOfficialJul 15, 2026Details
NVIDIA Nemotron 3 UltraOpenNVIDIA13.6%Nemotron 3 UltraOfficialJun 4, 2026Details
Qwen3.7 PlusAlibaba / Qwen10.2%Qwen 3.7 PlusOfficialDetails
SWE-1.6Cognition9.4%SWE-1.6OfficialApr 7, 2026Details
Mistral Medium 3.5Mistral AI8.0%Mistral 3.5 MediumOfficialApr 28, 2026Details

Each row reports the model’s weighted score (Main) on FrontierCode. Click a row for the full run context.