METR Task-Completion Time Horizons
Measures the length of software/ML-engineering tasks (in human-expert minutes) that an AI agent can complete with 50% reliability, derived from a logistic fit over HCAST, RE-Bench, and SWAA task suites.
What this benchmark measures
Measures the length of software/ML-engineering tasks (in human-expert minutes) that an AI agent can complete with 50% reliability, derived from a logistic fit over HCAST, RE-Bench, and SWAA task suites.
Rows on this page are sourced from public benchmark artifacts, leaderboard exports, or source-linked model reports. Each row keeps benchmark version, source model name, and available run details attached to the score.
The metric shown here is 50% time horizon. It should be interpreted within METR Task-Completion Time Horizons, not compared as part of a site-wide ranking.
Frequently asked
What is METR Task-Completion Time Horizons?
Measures the length of software/ML-engineering tasks (in human-expert minutes) that an AI agent can complete with 50% reliability, derived from a logistic fit over HCAST, RE-Bench, and SWAA task suites. It is a agents benchmark measured by 50% time horizon.
What does 50% time horizon mean on METR Task-Completion Time Horizons?
METR Task-Completion Time Horizons reports 50% time horizon (min); higher is better. Scores are shown only within METR Task-Completion Time Horizons and are never averaged with other benchmarks.
What is the top reported METR Task-Completion Time Horizons score?
Claude Mythos Preview has the top reported score on METR Task-Completion Time Horizons: 1044.8 min (50% time horizon).
Why do METR Task-Completion Time Horizons scores differ across runs?
Harness, scaffold, reasoning effort, and prompt setup change results, so two runs of the same model can differ. evals.report keeps each score with its run context so the differences stay visible.
Does evals.report rank models across benchmarks?
No. METR Task-Completion Time Horizons scores are shown within their own metric; evals.report never combines benchmarks into a composite ranking or a single "best model".