evals.report
BenchmarksLabsCompareRun guidesIn the wild

DeepSWE

A long-horizon software-engineering benchmark with original tasks, broad repository coverage, and behavioral verifiers.

Coding% resolvedHigher is better

What is DeepSWE?

A long-horizon software-engineering benchmark with original tasks, broad repository coverage, and behavioral verifiers. evals.report tracks reported DeepSWE scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.

Top reported DeepSWE score: GPT-5.6 Sol 73.0% (% resolved).

Community runs: 3 independent reproductions are listed separately from official scores, each shown with its own source and caveats — never merged or averaged. See how scores are labeled.

ModelLabScoreSource modelStatusDate
GPT-5.6 SolOpenAI73.0%GPT-5.6 SolUnverifiedJul 9, 2026Details
GPT-5.5OpenAI70.05%gpt-5-5OfficialApr 23, 2026Details
Claude Fable 5Anthropic70.0%Claude Fable 5UnverifiedJun 9, 2026Details
Claude Opus 5Anthropic68.8%Claude Opus 5VerifiedJul 24, 2026Details
Kimi K3OpenMoonshot AI67.5%Kimi K3VerifiedJul 16, 2026Details
GLM-5.3OpenZ.ai66.9%GLM-5.3VerifiedAug 14, 2026Details
Gemini 3.7 FlashGoogle DeepMind65.3%Gemini 3.7 FlashVerifiedAug 13, 2026Details
DeepSeek V4 Pro 0813OpenDeepSeek62.7%DeepSeek-V4-Pro-0813VerifiedAug 13, 2026Details
Claude Opus 4.8Anthropic58%Claude Opus 4.8 [max]VerifiedMay 28, 2026Details
GPT-5.4OpenAI55.53%gpt-5-4OfficialMar 5, 2026Details
DeepSeek V4 Flash 0731OpenDeepSeek54.4%DeepSeek-V4-Flash-0731VerifiedJul 31, 2026Details
Claude Opus 4.7Anthropic54.20%claude-opus-4-7OfficialApr 16, 2026Details
Gemini 3.6 FlashGoogle DeepMind48.6%Gemini 3.6 FlashVerifiedJul 21, 2026Details
GLM-5.2OpenZ.ai46.2%GLM-5.2VerifiedJun 16, 2026Details
Nex-N2-ProOpenNex AGI33.6%Nex-N2-ProVerifiedJun 2, 2026Details
Claude Sonnet 4.6Anthropic31.56%claude-sonnet-4-6OfficialFeb 17, 2026Details
Gemini 3.5 FlashGoogle DeepMind28.32%gemini-3-5-flashOfficialMay 19, 2026Details
Claude Opus 4.6Anthropic27.06%claude-opus-4-6OfficialFeb 5, 2026Details
Kimi K2.6OpenMoonshot AI23.89%kimi-k2-6OfficialApr 20, 2026Details
GLM-5.1OpenZ.ai17.48%glm-5-1OfficialApr 7, 2026Details
MiniMax M3OpenMiniMax13.3%MiniMax-M3 [default]CommunityJun 1, 2026Details
Gemini 3.1 Pro PreviewGoogle DeepMind9.88%gemini-3-1-pro-previewOfficialFeb 19, 2026Details
Nex-N2-miniOpenNex AGI8.0%Nex-N2-miniVerifiedJun 2, 2026Details
DeepSeek V4 ProOpenDeepSeek7.52%deepseek-v4-proOfficialApr 24, 2026Details
Community run · @ivanfioravanti (X)5.3%2.2DeepSeek V4 Pro [reasoning max]CommunityApr 24, 2026Details
Gemini 3 FlashGoogle DeepMind5.16%gemini-3-flash-previewOfficialDec 17, 2025Details
Qwen 3.6 PlusAlibaba / Qwen2.65%qwen3-6-plusOfficialApr 2, 2026Details
Qwen 3.6 27BOpenAlibaba / Qwen1.79%Qwen 3.6 27B (FP8)CommunityApr 22, 2026Details
Claude Haiku 4.5Anthropic0.22%claude-haiku-4-5OfficialOct 15, 2025Details

Each row reports the model’s % resolved on DeepSWE. Click a row for the full run context.