Terminal-Bench 2.0
An agentic benchmark measuring whether an AI model can complete real command-line / terminal software tasks end-to-end (version 2.0, the 89-task set), scored by task success rate. Distinct from the newer Terminal-Bench 2.1 (a different task set); most 2026 model cards self-report this 2.0 version.
What is Terminal-Bench 2.0?
An agentic benchmark measuring whether an AI model can complete real command-line / terminal software tasks end-to-end (version 2.0, the 89-task set), scored by task success rate. Distinct from the newer Terminal-Bench 2.1 (a different task set); most 2026 model cards self-report this 2.0 version. evals.report tracks reported Terminal-Bench 2.0 scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.
Top reported Terminal-Bench 2.0 score: Claude Fable 5 — 84.3% (task success).
| Model | Lab | Score↓ | Source model | Status | Date | |
|---|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | 84.3% | Claude Fable 5 | Verified | Jun 9, 2026 | Details |
| Claude Mythos Preview | Anthropic | 82.0% | — | Unverified | Apr 7, 2026 | Details |
| GPT-5.3-Codex | OpenAI | 77.3% | — | Verified | Feb 5, 2026 | Details |
| GPT-5.4 | OpenAI | 75.1% | — | Verified | Mar 5, 2026 | Details |
| Qwen3.7 Max Preview | Alibaba / Qwen | 69.7% | — | Unverified | May 14, 2026 | Details |
| MiMo-V2.5-ProOpen | Xiaomi | 68.4% | — | Verified | Apr 22, 2026 | Details |
| DeepSeek V4 ProOpen | DeepSeek | 67.9% | — | Verified | Apr 24, 2026 | Details |
| Kimi K2.6Open | Moonshot AI | 66.7% | — | Verified | Apr 20, 2026 | Details |
| Claude Opus 4.6 | Anthropic | 65.4% | — | Verified | Feb 5, 2026 | Details |
| GLM-5.1Open | Z.ai | 63.5% | — | Verified | Apr 7, 2026 | Details |
| Claude Sonnet 4.6 | Anthropic | 59.1% | — | Verified | Feb 17, 2026 | Details |
| MiniMax M2.7Open | MiniMax | 57.0% | — | Verified | Mar 18, 2026 | Details |
| DeepSeek V4 FlashOpen | DeepSeek | 56.9% | — | Verified | Apr 24, 2026 | Details |
| Doubao Seed 2.0 Pro | ByteDance | 55.8% | — | Verified | Feb 14, 2026 | Details |
| Qwen3.5-397B-A17BOpen | Alibaba / Qwen | 52.5% | — | Verified | Feb 16, 2026 | Details |
| MAI-Thinking-1 | Microsoft AI | 46.0% | — | Verified | Jun 2, 2026 | Details |
| GLM-4.7Open | Z.ai | 41.0% | — | Verified | Dec 22, 2025 | Details |
| DeepSeek V3.2Open | DeepSeek | 39.6% | — | Official | Dec 1, 2025 | Details |
| Kimi K2 ThinkingOpen | Moonshot AI | 35.7% | — | Official | Nov 6, 2025 | Details |
Each row reports the model’s task success on Terminal-Bench 2.0. Click a row for the full run context.