evals.report
BenchmarksLabsCompareRun guidesIn the wild
BenchmarksReasoning

EnigmaEval

A benchmark of 1,184 puzzle-hunt challenges spanning text and images that probes models' ability to perform implicit knowledge synthesis, lateral thinking, and multi-step deductive reasoning to uncover hidden solution paths.

ReasoningaccuracyHigher is better

What this benchmark measures

A benchmark of 1,184 puzzle-hunt challenges spanning text and images that probes models' ability to perform implicit knowledge synthesis, lateral thinking, and multi-step deductive reasoning to uncover hidden solution paths.

Rows on this page are sourced from public benchmark artifacts, leaderboard exports, or source-linked model reports. Each row keeps benchmark version, source model name, and available run details attached to the score.

The metric shown here is accuracy. It should be interpreted within EnigmaEval, not compared as part of a site-wide ranking.

No composite ranking
evals.report never combines benchmarks. accuracy on EnigmaEval is its own number — don’t average it with other metrics.

Frequently asked

What is EnigmaEval?

A benchmark of 1,184 puzzle-hunt challenges spanning text and images that probes models' ability to perform implicit knowledge synthesis, lateral thinking, and multi-step deductive reasoning to uncover hidden solution paths. It is a reasoning benchmark measured by accuracy.

What does accuracy mean on EnigmaEval?

EnigmaEval reports accuracy (%); higher is better. Scores are shown only within EnigmaEval and are never averaged with other benchmarks.

What is the top reported EnigmaEval score?

GPT-5.4 Pro has the top reported score on EnigmaEval: 23.82% (accuracy).

Why do EnigmaEval scores differ across runs?

Harness, scaffold, reasoning effort, and prompt setup change results, so two runs of the same model can differ. evals.report keeps each score with its run context so the differences stay visible.

Does evals.report rank models across benchmarks?

No. EnigmaEval scores are shown within their own metric; evals.report never combines benchmarks into a composite ranking or a single "best model".