evals.report
BenchmarksLabsCompareRun guidesIn the wild
BenchmarksReasoning

PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts

A multilingual mathematical reasoning benchmark of 9,000 parallel problems across 18 languages and 4 difficulty levels (K-12 to Olympiad/frontier), scored by difficulty-weighted accuracy.

ReasoningDifficulty-Weighted Accuracy (DW-ACC)Higher is better

What this benchmark measures

A multilingual mathematical reasoning benchmark of 9,000 parallel problems across 18 languages and 4 difficulty levels (K-12 to Olympiad/frontier), scored by difficulty-weighted accuracy.

Rows on this page are sourced from public benchmark artifacts, leaderboard exports, or source-linked model reports. Each row keeps benchmark version, source model name, and available run details attached to the score.

The metric shown here is Difficulty-Weighted Accuracy (DW-ACC). It should be interpreted within PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts, not compared as part of a site-wide ranking.

No composite ranking
evals.report never combines benchmarks. Difficulty-Weighted Accuracy (DW-ACC) on PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts is its own number — don’t average it with other metrics.

Frequently asked

What is PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts?

A multilingual mathematical reasoning benchmark of 9,000 parallel problems across 18 languages and 4 difficulty levels (K-12 to Olympiad/frontier), scored by difficulty-weighted accuracy. It is a reasoning benchmark measured by Difficulty-Weighted Accuracy (DW-ACC).

What does Difficulty-Weighted Accuracy (DW-ACC) mean on PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts?

PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts reports Difficulty-Weighted Accuracy (DW-ACC) (%); higher is better. Scores are shown only within PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts and are never averaged with other benchmarks.

What is the top reported PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts score?

Kimi K2 Instruct has the top reported score on PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts: 65.1% (Difficulty-Weighted Accuracy (DW-ACC)).

Why do PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts scores differ across runs?

Harness, scaffold, reasoning effort, and prompt setup change results, so two runs of the same model can differ. evals.report keeps each score with its run context so the differences stay visible.

Does evals.report rank models across benchmarks?

No. PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts scores are shown within their own metric; evals.report never combines benchmarks into a composite ranking or a single "best model".