evals.report
BenchmarksLabsCompareRun guidesIn the wild
BenchmarksReasoning

PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts

A multilingual mathematical reasoning benchmark of 9,000 parallel problems across 18 languages and 4 difficulty levels (K-12 to Olympiad/frontier), scored by difficulty-weighted accuracy.

ReasoningDifficulty-Weighted Accuracy (DW-ACC)Higher is better

What is PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts?

A multilingual mathematical reasoning benchmark of 9,000 parallel problems across 18 languages and 4 difficulty levels (K-12 to Olympiad/frontier), scored by difficulty-weighted accuracy. evals.report tracks reported PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts scores with the model, source, status, date, and run caveats attached — official leaderboard scores, vendor-reported launches, and clearly labeled community runs.

Top reported PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts score: Kimi K2 Instruct 65.1% (Difficulty-Weighted Accuracy (DW-ACC)).

ModelLabScoreSource modelStatusDate
Kimi K2 InstructOpenMoonshot AI65.1%UnverifiedJul 11, 2025Details
Gemini 2.5 ProGoogle DeepMind52.2VerifiedMar 25, 2025Details
DeepSeek R1OpenDeepSeek47.0VerifiedJan 20, 2025Details
o4-miniOpenAI45.6VerifiedApr 16, 2025Details
Claude 3.7 SonnetAnthropic33.5VerifiedFeb 24, 2025Details
DeepSeek V3 0324OpenDeepSeek30.7VerifiedMar 24, 2025Details
GPT-4.1OpenAI26.4VerifiedApr 14, 2025Details
Llama 4 MaverickOpenMeta26.1VerifiedApr 5, 2025Details
Llama 4 ScoutOpenMeta20.9VerifiedApr 5, 2025Details
DeepSeek V3OpenDeepSeek20.4VerifiedDec 26, 2024Details
GPT-4oOpenAI13.7VerifiedMay 13, 2024Details

Each row reports the model’s Difficulty-Weighted Accuracy (DW-ACC) on PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts. Click a row for the full run context.