evals.report
BenchmarksLabsCompareRun guidesIn the wild

Terminal-Bench 2.0

An agentic benchmark measuring whether an AI model can complete real command-line / terminal software tasks end-to-end (version 2.0, the 89-task set), scored by task success rate. Distinct from the newer Terminal-Bench 2.1 (a different task set); most 2026 model cards self-report this 2.0 version.

Agentstask successHigher is better

What is Terminal-Bench 2.0?

An agentic benchmark measuring whether an AI model can complete real command-line / terminal software tasks end-to-end (version 2.0, the 89-task set), scored by task success rate. Distinct from the newer Terminal-Bench 2.1 (a different task set); most 2026 model cards self-report this 2.0 version. evals.report tracks reported Terminal-Bench 2.0 scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.

Top reported Terminal-Bench 2.0 score: Claude Fable 5 84.3% (task success).

ModelLabScoreSource modelStatusDate
Claude Fable 5Anthropic84.3%Claude Fable 5VerifiedJun 9, 2026Details
Claude Mythos PreviewAnthropic82.0%UnverifiedApr 7, 2026Details
GPT-5.3-CodexOpenAI77.3%VerifiedFeb 5, 2026Details
GPT-5.4OpenAI75.1%VerifiedMar 5, 2026Details
Qwen3.7 Max PreviewAlibaba / Qwen69.7%UnverifiedMay 14, 2026Details
MiMo-V2.5-ProOpenXiaomi68.4%VerifiedApr 22, 2026Details
DeepSeek V4 ProOpenDeepSeek67.9%VerifiedApr 24, 2026Details
Kimi K2.6OpenMoonshot AI66.7%VerifiedApr 20, 2026Details
Claude Opus 4.6Anthropic65.4%VerifiedFeb 5, 2026Details
GLM-5.1OpenZ.ai63.5%VerifiedApr 7, 2026Details
Claude Sonnet 4.6Anthropic59.1%VerifiedFeb 17, 2026Details
MiniMax M2.7OpenMiniMax57.0%VerifiedMar 18, 2026Details
DeepSeek V4 FlashOpenDeepSeek56.9%VerifiedApr 24, 2026Details
Doubao Seed 2.0 ProByteDance55.8%VerifiedFeb 14, 2026Details
Qwen3.5-397B-A17BOpenAlibaba / Qwen52.5%VerifiedFeb 16, 2026Details
MAI-Thinking-1Microsoft AI46.0%VerifiedJun 2, 2026Details
GLM-4.7OpenZ.ai41.0%VerifiedDec 22, 2025Details
DeepSeek V3.2OpenDeepSeek39.6%OfficialDec 1, 2025Details
Kimi K2 ThinkingOpenMoonshot AI35.7%OfficialNov 6, 2025Details

Each row reports the model’s task success on Terminal-Bench 2.0. Click a row for the full run context.