evals.report
BenchmarksLabsCompareRun guidesIn the wild

FrontierCode

Cognition's benchmark for code mergeability and production quality, not just correctness. Tasks are drawn from 36 real open-source repositories and authored by their maintainers (40+ hours each), with concise, humanlike prompts (~1/3 the length of SWE-bench Pro). Solutions are graded against a maintainer-style rubric spanning behavioral correctness, regression safety, mechanical cleanliness, test correctness, scope, and code quality; the reported score is a weighted aggregate of the rubric items, and any solution that fails a 'blocker' criterion scores 0. Revision 1.1 (2026-07-07) publishes two nested subsets — Main (100 tasks) and Extended (150) — having deprecated the original Diamond (50 hardest) subset, and zeroes runs flagged for consulting solution-bearing sources such as the original pull request. Each model is run 5× at every available reasoning effort and its best effort is reported. Tasks are kept private to avoid contamination.

Codingweighted score (Main)Higher is better

What is FrontierCode?

Cognition's benchmark for code mergeability and production quality, not just correctness. Tasks are drawn from 36 real open-source repositories and authored by their maintainers (40+ hours each), with concise, humanlike prompts (~1/3 the length of SWE-bench Pro). Solutions are graded against a maintainer-style rubric spanning behavioral correctness, regression safety, mechanical cleanliness, test correctness, scope, and code quality; the reported score is a weighted aggregate of the rubric items, and any solution that fails a 'blocker' criterion scores 0. Revision 1.1 (2026-07-07) publishes two nested subsets — Main (100 tasks) and Extended (150) — having deprecated the original Diamond (50 hardest) subset, and zeroes runs flagged for consulting solution-bearing sources such as the original pull request. Each model is run 5× at every available reasoning effort and its best effort is reported. Tasks are kept private to avoid contamination. evals.report tracks reported FrontierCode scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.

Top reported FrontierCode score: Claude Fable 5 53.5% (weighted score (Main)).

ModelLabScoreSource modelStatusDate
Claude Fable 5Anthropic53.5%Fable 5OfficialJun 9, 2026Details
Claude Opus 5Anthropic53.4%Opus 5OfficialJul 24, 2026Details
GPT-5.6 SolOpenAI47.5%GPT-5.6 SolOfficialJul 9, 2026Details
Claude Opus 4.8Anthropic46.5%Opus 4.8OfficialMay 28, 2026Details
GPT-5.5OpenAI43.0%GPT-5.5OfficialApr 23, 2026Details
Claude Sonnet 5Anthropic42.7%Sonnet 5OfficialJun 30, 2026Details
Grok 4.5xAI42.4%Grok 4.5OfficialJul 8, 2026Details
SWE-1.7Cognition42.3%SWE-1.7OfficialJul 8, 2026Details
GPT-5.6 TerraOpenAI41.3%GPT-5.6 TerraOfficialJul 9, 2026Details
GPT-5.6 LunaOpenAI39.8%GPT-5.6 LunaOfficialJul 9, 2026Details
Claude Opus 4.7Anthropic38.5%Opus 4.7OfficialApr 16, 2026Details
GLM-5.2OpenZ.ai24.5%GLM 5.2OfficialJun 16, 2026Details
DeepSeek V4 ProOpenDeepSeek17.6%DeepSeek V4 ProOfficialApr 24, 2026Details
MiniMax M3OpenMiniMax14.7%MiniMax M3OfficialJun 1, 2026Details
InklingOpenThinking Machines Lab14.0%InklingOfficialJul 15, 2026Details
SWE-1.6Cognition9.4%SWE-1.6OfficialApr 7, 2026Details

Each row reports the model’s weighted score (Main) on FrontierCode. Click a row for the full run context.