LongBench v2
A long-context benchmark of 503 challenging multiple-choice questions with contexts from 8k to 2M words across six task categories, designed to test deep understanding and reasoning over realistic long-context multitasks.
What this benchmark measures
A long-context benchmark of 503 challenging multiple-choice questions with contexts from 8k to 2M words across six task categories, designed to test deep understanding and reasoning over realistic long-context multitasks.
Rows on this page are sourced from public benchmark artifacts, leaderboard exports, or source-linked model reports. Each row keeps benchmark version, source model name, and available run details attached to the score.
The metric shown here is accuracy. It should be interpreted within LongBench v2, not compared as part of a site-wide ranking.
Frequently asked
What is LongBench v2?
A long-context benchmark of 503 challenging multiple-choice questions with contexts from 8k to 2M words across six task categories, designed to test deep understanding and reasoning over realistic long-context multitasks. It is a reasoning benchmark measured by accuracy.
What does accuracy mean on LongBench v2?
LongBench v2 reports accuracy (%); higher is better. Scores are shown only within LongBench v2 and are never averaged with other benchmarks.
What is the top reported LongBench v2 score?
Gemini 2.5 Pro has the top reported score on LongBench v2: 63.3% (accuracy).
Why do LongBench v2 scores differ across runs?
Harness, scaffold, reasoning effort, and prompt setup change results, so two runs of the same model can differ. evals.report keeps each score with its run context so the differences stay visible.
Does evals.report rank models across benchmarks?
No. LongBench v2 scores are shown within their own metric; evals.report never combines benchmarks into a composite ranking or a single "best model".