GPT-5.5 vs Claude Opus 4.8
33 shared benchmarks
13
wins
0
ties
20
wins
38.5%
APEX-Agents
39.4%
0.43%
ARC-AGI 3
1.5%
46.9%
Agents' Last Exam
45.2%
22.7%
AutomationBench
27.2%
83.6%
BabyVision
81.2%
84.4%
BrowseComp
84.3%
84.1%
CharXiv
80.5%
67.0%
DeepSWE
59.0%
47.9%
ExploitBench
40.0%
51.7%
FrontierMath
80.0%
35.4%
FrontierMath Tier 4
56.1%
64.9%
FrontierSWE
66.7%
93.5%
GPQA Diamond
91.0%
41.4%
Humanity's Last Exam
49.8%
38.3%
JobBench
48.4%
92.2%
MATH-Vision
86.7%
82.8%
MCP-Atlas
77.8%
35.5%
MLS Bench
42.8%
81.2%
MMMU Pro
78.9%
49.5%
OSWorld 2.0
54.8%
60.9%
OfficeQA Pro
63.9%
89.4%
OmniDocBench
87.9%
55.8%
PerceptionBench
47.2%
28.4%
PostTrain Bench
34.1%
70.8%
Program Bench
71.9%
14.0%
SWE Marathon
40.0%
58.6%
SWE-bench Pro
69.2%
29.05%
SpreadsheetBench 2
31.55%
83.4%
Terminal-Bench 2.1
84.6%
55.6%
Toolathlon
59.9%
73.5%
Toolathlon-Verified
76.2%
38.5%
WorldVQA
39.1%
22.0%
ZeroBench
17.0%