FrontierCode
Cognition's benchmark for code mergeability and production quality, not just correctness. Tasks are drawn from 36 real open-source repositories and authored by their maintainers (40+ hours each), with concise, humanlike prompts (~1/3 the length of SWE-bench Pro). Solutions are graded against a maintainer-style rubric spanning behavioral correctness, regression safety, mechanical cleanliness, test correctness, scope, and code quality; the reported score is a weighted aggregate of the rubric items, and any solution that fails a 'blocker' criterion scores 0. Revision 1.1 (2026-07-07) publishes two nested subsets — Main (100 tasks) and Extended (150) — having deprecated the original Diamond (50 hardest) subset, and zeroes runs flagged for consulting solution-bearing sources such as the original pull request. Each model is run 5× at every available reasoning effort and its best effort is reported. Tasks are kept private to avoid contamination.
No run guide for this benchmark yet.