evals.report
BenchmarksLabsCompareRun guidesIn the wild

BrowseComp

A benchmark of 1,266 hard-to-find, multi-hop web-browsing questions whose answers are difficult to locate but easy to verify, measuring an agent's ability to persistently search and synthesize information from the web.

AgentsaccuracyHigher is better

What is BrowseComp?

A benchmark of 1,266 hard-to-find, multi-hop web-browsing questions whose answers are difficult to locate but easy to verify, measuring an agent's ability to persistently search and synthesize information from the web. evals.report tracks reported BrowseComp scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.

Top reported BrowseComp score: GPT-5.6 Sol Ultra 92.2% (accuracy).

ModelLabScoreSource modelStatusDate
GPT-5.6 Sol UltraOpenAI92.2%GPT-5.6 Sol (Ultra)VerifiedJul 9, 2026Details
Kimi K3OpenMoonshot AI91.2%Kimi K3VerifiedJul 17, 2026Details
GPT-5.6 SolOpenAI90.4%GPT-5.6 SolVerifiedJul 9, 2026Details
GPT-5.6 TerraOpenAI87.5%GPT-5.6 TerraVerifiedJul 9, 2026Details
Nex-N2-ProOpenNex AGI83.7%Nex-N2-ProVerifiedJun 2, 2026Details
GPT-5.6 LunaOpenAI83.3%GPT-5.6 LunaVerifiedJul 9, 2026Details
InklingOpenThinking Machines Lab77.1%InklingVerifiedJul 15, 2026Details
Nex-N2-miniOpenNex AGI74.1%Nex-N2-miniVerifiedJun 2, 2026Details
Kimi K2 ThinkingOpenMoonshot AI60.2%VerifiedNov 6, 2025Details
GPT-5OpenAI54.9%VerifiedAug 7, 2025Details
o3OpenAI49.7%VerifiedApr 16, 2025Details
NVIDIA Nemotron 3 UltraOpenNVIDIA44.4%Nemotron-3-Ultra-550B-A55B (BF16)VerifiedJun 4, 2026Details
DeepSeek V3.2OpenDeepSeek40.1%UnverifiedDec 1, 2025Details
o4-miniOpenAI28.3%VerifiedApr 16, 2025Details
Claude Sonnet 4.5Anthropic24.1%UnverifiedSep 29, 2025Details
GPT-4oOpenAI0.6%VerifiedMay 13, 2024Details

Each row reports the model’s accuracy on BrowseComp. Click a row for the full run context.