benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
Z.ai
GLM-5.2
27 benchmarks
AIME
#4 of 117
99.2%
HMMT Nov 2025
#3 of 4
94.4%
HMMT Feb 2026
#4 of 5
92.5%
GPQA Diamond
#23 of 118
91.2%
Terminal-Bench 2.1
#13 of 27
82.7%
MCP-Atlas
#7 of 17
82.6%
Terminal-Bench 2.1
#13 of 27
81.0%
SWE-bench Verified
#11 of 86
80.0%
MCP-Atlas
#7 of 17
76.8%
FrontierSWE
#4 of 8
74.4%
FrontierSWE
#4 of 8
67.3%
Program Bench
#6 of 6
63.7%
SWE-bench Pro
#10 of 43
62.1%
Toolathlon-Verified
#6 of 7
59.9%
Humanity's Last Exam
#8 of 72
54.7%
Toolathlon
#7 of 7
48.2%
DeepSWE
#16 of 22
46.2%
DeepSWE
#16 of 22
44.0%
JobBench
#5 of 6
43.4%
OfficeQA Pro
#6 of 6
41.4%
Humanity's Last Exam
#8 of 72
40.5%
MLS Bench
#5 of 6
40.4%
APEX-Agents
#7 of 7
35.6%
PostTrain Bench
#4 of 6
34.3%
SpreadsheetBench 2
#6 of 6
28.12%
SWE Marathon
#6 of 6
13.0%
AutomationBench
#6 of 6
12.9%