LongBench v2
A long-context benchmark of 503 challenging multiple-choice questions with contexts from 8k to 2M words across six task categories, designed to test deep understanding and reasoning over realistic long-context multitasks.
What is LongBench v2?
A long-context benchmark of 503 challenging multiple-choice questions with contexts from 8k to 2M words across six task categories, designed to test deep understanding and reasoning over realistic long-context multitasks. evals.report tracks reported LongBench v2 scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.
Top reported LongBench v2 score: Gemini 2.5 Pro — 63.3% (accuracy).
| Model | Lab | Score↓ | Source model | Status | Date | |
|---|---|---|---|---|---|---|
| Gemini 2.5 Pro | Google DeepMind | 63.3% | — | Official | Mar 25, 2025 | Details |
| Gemini 2.5 Flash | Google DeepMind | 62.1% | — | Official | Apr 17, 2025 | Details |
| Qwen3 235B A22B Instruct 2507Open | Alibaba / Qwen | 58.3% | — | Official | Jul 21, 2025 | Details |
| DeepSeek R1Open | DeepSeek | 58.3% | — | Official | Jan 20, 2025 | Details |
| GPT-4o | OpenAI | 51.4% | — | Official | May 13, 2024 | Details |
| Gemini 2.0 Flash | Google DeepMind | 51.1% | — | Official | Dec 11, 2024 | Details |
| Claude 3.5 Sonnet | Anthropic | 46.7% | — | Official | Jun 20, 2024 | Details |
| Kimi K2 InstructOpen | Moonshot AI | 44.3% | — | Official | Jul 11, 2025 | Details |
| Mistral Large | Mistral AI | 39.6% | — | Official | Feb 26, 2024 | Details |
Each row reports the model’s accuracy on LongBench v2. Click a row for the full run context.