benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
head to head
GPT-5.5
vs
GLM-5.2
18 shared benchmarks
13
wins
0
ties
5
wins
38.5%
APEX-Agents
35.6%
22.7%
AutomationBench
12.9%
67.0%
DeepSWE
44.0%
64.9%
FrontierSWE
67.3%
93.5%
GPQA Diamond
91.2%
41.4%
Humanity's Last Exam
40.5%
38.3%
JobBench
43.4%
82.8%
MCP-Atlas
76.8%
35.5%
MLS Bench
40.4%
60.9%
OfficeQA Pro
41.4%
28.4%
PostTrain Bench
34.3%
70.8%
Program Bench
63.7%
14.0%
SWE Marathon
13.0%
58.6%
SWE-bench Pro
62.1%
29.05%
SpreadsheetBench 2
28.12%
83.4%
Terminal-Bench 2.1
81.0%
55.6%
Toolathlon
48.2%
73.5%
Toolathlon-Verified
59.9%