evals.report
BenchmarksLabsCompareRun guidesIn the wild

GDPval

GDPval evaluates AI models agentically (shell + web access via a sandbox harness) on real-world economically valuable knowledge-work deliverables — documents, spreadsheets, slides, diagrams — spanning 44 occupations across 9 major U.S. GDP industries, scored by blind pairwise quality comparison; the Artificial Analysis GDPval-AA variant reports results as an Elo rating.

AgentsEloHigher is better

What this benchmark measures

GDPval evaluates AI models agentically (shell + web access via a sandbox harness) on real-world economically valuable knowledge-work deliverables — documents, spreadsheets, slides, diagrams — spanning 44 occupations across 9 major U.S. GDP industries, scored by blind pairwise quality comparison; the Artificial Analysis GDPval-AA variant reports results as an Elo rating.

Rows on this page are sourced from public benchmark artifacts, leaderboard exports, or source-linked model reports. Each row keeps benchmark version, source model name, and available run details attached to the score.

The metric shown here is Elo. It should be interpreted within GDPval, not compared as part of a site-wide ranking.

No composite ranking
evals.report never combines benchmarks. Elo on GDPval is its own number — don’t average it with other metrics.

Frequently asked

What is GDPval?

GDPval evaluates AI models agentically (shell + web access via a sandbox harness) on real-world economically valuable knowledge-work deliverables — documents, spreadsheets, slides, diagrams — spanning 44 occupations across 9 major U.S. GDP industries, scored by blind pairwise quality comparison; the Artificial Analysis GDPval-AA variant reports results as an Elo rating. It is a agents benchmark measured by Elo.

What does Elo mean on GDPval?

GDPval reports Elo; higher is better. Scores are shown only within GDPval and are never averaged with other benchmarks.

What is the top reported GDPval score?

Claude Fable 5 has the top reported score on GDPval: 1932 (Elo).

Why do GDPval scores differ across runs?

Harness, scaffold, reasoning effort, and prompt setup change results, so two runs of the same model can differ. evals.report keeps each score with its run context so the differences stay visible.

Does evals.report rank models across benchmarks?

No. GDPval scores are shown within their own metric; evals.report never combines benchmarks into a composite ranking or a single "best model".