evals.report
BenchmarksLabsCompareRun guidesIn the wild
Z.aiGLM

GLM-5.2

Z.ai · GLM. Released Jun 16, 2026.

GLM-5.2 is a model from Z.ai in the GLM family, released Jun 16, 2026. evals.report tracks 14 reported GLM-5.2 benchmark scores across ARC-AGI-1, ARC-AGI-2, Humanity's Last Exam, AIME 2026, GPQA Diamond, MathArena HMMT February 2026, SWE-bench Pro, DeepSWE, and 6 more — each shown with its benchmark, metric, source status, and date, and never combined into a single ranking.

Open14 results

Benchmark results 14

Compare this model
BenchmarkCategoryScoreMetricStatusDate
ARC-AGI-1Reasoning77%accuracyOfficialJun 16, 2026Details
ARC-AGI-2Reasoning22.78%accuracyOfficialJun 16, 2026Details
Humanity's Last ExamReasoning40.5%accuracyVerifiedJun 16, 2026Details
AIME 2026Reasoning99.2%accuracyVerifiedJun 16, 2026Details
GPQA DiamondReasoning91.2%accuracyVerifiedJun 16, 2026Details
MathArena HMMT February 2026Reasoning92.5%accuracyVerifiedJun 16, 2026Details
SWE-bench ProCoding62.1%% resolvedVerifiedJun 16, 2026Details
DeepSWECoding46.2%% resolvedVerifiedJun 16, 2026Details
Terminal-Bench 2.1Agents81.0%task successVerifiedJun 16, 2026Details
MCP AtlasTool use76.8%pass rateVerifiedJun 16, 2026Details
SWE-MarathonAgents13.0%resolution rate (pass@1)VerifiedJun 16, 2026Details
FrontierSWEAgents74%dominance scoreOfficialJun 16, 2026Details
PostTrainBenchAgents34.3%weighted average scoreVerifiedJun 16, 2026Details
FrontierCodeCoding24.5%weighted score (Main)OfficialJun 16, 2026Details

In the wild 5

See all

Real-world feedback on GLM-5.2 from people using it on actual prompts — praise and criticism alike, each linked to its source. Qualitative, never scored.

Report tone

Report type

Topic

Pranav Sriram
X·@PranavSriram1·
Negative
For my research, Fable felt like a clear step change … I was excited about the GLM 5.2 hype and tried it; sadly it's nowhere close

Task Evaluating models for research work (alongside Fable and GPT-5.5 Pro).

anecdotal
@Hesamation
X·@Hesamation·
Positive
GLM 5.2 ranks unusually high on FrontierSWE (long-horizon agentic engineering) … using it with OpenCode is also not far from the quality of Claude Code or Codex.

Task Day-to-day agentic coding with GLM-5.2 in OpenCode.

anecdotalhigh-signal user
Benchmark reproductionField testCodingAgents
View on X
Guillermo Rauch
X·@rauchg·
Positive
Genuinely impressed, almost shocked, at how good GLM-5.2 … is at coding. This changes things.
anecdotalhigh-signal user
Theo
X·@theo·
Mixed
Having an open weight model surpass GPT-5.4 and every Gemini model is dope. That said - it's not cheap. Both Opus 4.8 and GPT-5.5 set to "medium" are cheaper and smarter than GLM-5.2
anecdotalhigh-signal user
Hassan
X·@nutlope·
Positive
Asked GLM 5.2 (left) and Opus 4.8 (right) to design a menu and GLM has even better taste while being 4x cheaper. GLM also nailed a lot of small details like adding "chef's pick" and "vegetatian" tags to some dishes.

Task Asked GLM-5.2 and Opus 4.8 to design a restaurant menu (UI/design).

high-signal useroutput shown
Field testComparisonCostCodingMultimodal
View on X