AA-Omniscience: Knowledge and Hallucination Benchmark
A factuality and knowledge benchmark of 6,000 questions across 42 economically relevant topics in six domains, scoring models on the AA-Omniscience Index (-100 to 100) that rewards correct answers, penalizes hallucinations, and applies no penalty for abstaining.
What is AA-Omniscience: Knowledge and Hallucination Benchmark?
A factuality and knowledge benchmark of 6,000 questions across 42 economically relevant topics in six domains, scoring models on the AA-Omniscience Index (-100 to 100) that rewards correct answers, penalizes hallucinations, and applies no penalty for abstaining. evals.report tracks reported AA-Omniscience: Knowledge and Hallucination Benchmark scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.
Top reported AA-Omniscience: Knowledge and Hallucination Benchmark score: Gemini 3.1 Pro Preview — 33 (AA-Omniscience Index).
| Model | Lab | Score↓ | Source model | Status | Date | |
|---|---|---|---|---|---|---|
| Gemini 3.1 Pro Preview | Google DeepMind | 33 | — | Official | Feb 19, 2026 | Details |
| Claude Opus 4.8 | Anthropic | 27 | — | Official | May 28, 2026 | Details |
| Claude Opus 4.7 | Anthropic | 26 | — | Official | Apr 16, 2026 | Details |
| Gemini 3.5 Flash | Google DeepMind | 23 | — | Official | May 19, 2026 | Details |
| GPT-5.5 | OpenAI | 20 | — | Official | Apr 23, 2026 | Details |
| Grok 4.3 | xAI | 18 | — | Official | Apr 17, 2026 | Details |
| Claude Sonnet 4.6 | Anthropic | 12 | — | Official | Feb 17, 2026 | Details |
| Kimi K2.6Open | Moonshot AI | 6 | — | Official | Apr 20, 2026 | Details |
| GPT-5.4 | OpenAI | 6 | — | Official | Mar 5, 2026 | Details |
| Muse Spark | Meta | 4 | — | Official | Apr 8, 2026 | Details |
| MiMo-V2.5-ProOpen | Xiaomi | 4 | — | Official | Apr 22, 2026 | Details |
| InklingOpen | Thinking Machines Lab | 2.1 | Inkling | Verified | Jul 15, 2026 | Details |
| GLM-5.1Open | Z.ai | 2 | — | Official | Apr 7, 2026 | Details |
| MiniMax M2.7Open | MiniMax | 1 | — | Official | Mar 18, 2026 | Details |
| Claude Haiku 4.5 | Anthropic | -4 | — | Official | Oct 15, 2025 | Details |
| DeepSeek V4 ProOpen | DeepSeek | -10 | — | Official | Apr 24, 2026 | Details |
| Llama 3.1 405BOpen | Meta | -17 | — | Official | Jul 23, 2024 | Details |
| DeepSeek V4 FlashOpen | DeepSeek | -23 | — | Official | Apr 24, 2026 | Details |
| Qwen3.5-397B-A17BOpen | Alibaba / Qwen | -30 | — | Official | Feb 16, 2026 | Details |
| Mistral Medium 3.5 | Mistral AI | -36 | — | Official | Apr 28, 2026 | Details |
| NVIDIA Nemotron 3 Super 120B-A12BOpen | NVIDIA | -42 | — | Official | Mar 10, 2026 | Details |
| GPT-OSS-120BOpen | OpenAI | -50 | — | Official | Aug 5, 2025 | Details |
Each row reports the model’s AA-Omniscience Index on AA-Omniscience: Knowledge and Hallucination Benchmark. Click a row for the full run context.