BrowseComp
A benchmark of 1,266 hard-to-find, multi-hop web-browsing questions whose answers are difficult to locate but easy to verify, measuring an agent's ability to persistently search and synthesize information from the web.
What is BrowseComp?
A benchmark of 1,266 hard-to-find, multi-hop web-browsing questions whose answers are difficult to locate but easy to verify, measuring an agent's ability to persistently search and synthesize information from the web. evals.report tracks reported BrowseComp scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.
Top reported BrowseComp score: GPT-5.6 Sol Ultra — 92.2% (accuracy).
| Model | Lab | Score↓ | Source model | Status | Date | |
|---|---|---|---|---|---|---|
| GPT-5.6 Sol Ultra | OpenAI | 92.2% | GPT-5.6 Sol (Ultra) | Verified | Jul 9, 2026 | Details |
| Kimi K3Open | Moonshot AI | 91.2% | Kimi K3 | Verified | Jul 17, 2026 | Details |
| GPT-5.6 Sol | OpenAI | 90.4% | GPT-5.6 Sol | Verified | Jul 9, 2026 | Details |
| GPT-5.6 Terra | OpenAI | 87.5% | GPT-5.6 Terra | Verified | Jul 9, 2026 | Details |
| Nex-N2-ProOpen | Nex AGI | 83.7% | Nex-N2-Pro | Verified | Jun 2, 2026 | Details |
| GPT-5.6 Luna | OpenAI | 83.3% | GPT-5.6 Luna | Verified | Jul 9, 2026 | Details |
| InklingOpen | Thinking Machines Lab | 77.1% | Inkling | Verified | Jul 15, 2026 | Details |
| Nex-N2-miniOpen | Nex AGI | 74.1% | Nex-N2-mini | Verified | Jun 2, 2026 | Details |
| Kimi K2 ThinkingOpen | Moonshot AI | 60.2% | — | Verified | Nov 6, 2025 | Details |
| GPT-5 | OpenAI | 54.9% | — | Verified | Aug 7, 2025 | Details |
| o3 | OpenAI | 49.7% | — | Verified | Apr 16, 2025 | Details |
| NVIDIA Nemotron 3 UltraOpen | NVIDIA | 44.4% | Nemotron-3-Ultra-550B-A55B (BF16) | Verified | Jun 4, 2026 | Details |
| DeepSeek V3.2Open | DeepSeek | 40.1% | — | Unverified | Dec 1, 2025 | Details |
| o4-mini | OpenAI | 28.3% | — | Verified | Apr 16, 2025 | Details |
| Claude Sonnet 4.5 | Anthropic | 24.1% | — | Unverified | Sep 29, 2025 | Details |
| GPT-4o | OpenAI | 0.6% | — | Verified | May 13, 2024 | Details |
Each row reports the model’s accuracy on BrowseComp. Click a row for the full run context.