evals.report
BenchmarksLabsCompareRun guidesIn the wild

FrontierSWE

Proximal Labs' ultra-long-horizon coding-agent benchmark: 17 open-ended technical projects spanning implementation, performance engineering, and applied ML research (e.g. optimizing a real compiler, inventing better ML optimizers, building a PostgreSQL-compatible server backed by SQLite). Agents get up to 20 hours per task and 5 trials each; tasks are graded 0–1 on partial progress, and frontier models barely make headway — making FrontierSWE one of the few unsaturated public coding benchmarks. Models are ranked by 'dominance' (win rate against a random opponent across tasks).

Agentsdominance scoreHigher is better

Real-world posts that speak to FrontierSWE — qualitative feedback from people using these models, each linked to its source. Never scored or merged with the leaderboard above.

Report tone

Report type

Topic

@Hesamation · on GLM-5.2
X·@Hesamation·
Positive
GLM 5.2 ranks unusually high on FrontierSWE (long-horizon agentic engineering) … using it with OpenCode is also not far from the quality of Claude Code or Codex.

Task Day-to-day agentic coding with GLM-5.2 in OpenCode.

anecdotalhigh-signal user
Benchmark reproductionField testCodingAgents
View on X