Claude Opus 4.8 vs GPT-5.6 Sol
32 shared benchmarks
8
wins
1
ties
23
wins
39.4%
APEX-Agents
39.9%
1.5%
ARC-AGI 3
7.78%
45.2%
Agents' Last Exam
52.7%
27.2%
AutomationBench
29.7%
81.2%
BabyVision
88.9%
84.3%
BrowseComp
90.4%
80.5%
CharXiv
84.6%
59.0%
DeepSWE
72.7%
40.0%
ExploitBench
73.5%
80.0%
FrontierMath
83.0%
66.7%
FrontierSWE
71.3%
91.0%
GPQA Diamond
94.6%
49.8%
Humanity's Last Exam
44.5%
48.4%
JobBench
46.5%
86.7%
MATH-Vision
95.8%
77.8%
MCP-Atlas
83.6%
42.8%
MLS Bench
46.2%
78.9%
MMMU Pro
83.0%
54.8%
OSWorld 2.0
62.6%
63.9%
OfficeQA Pro
63.2%
87.9%
OmniDocBench
85.8%
47.2%
PerceptionBench
59.7%
34.1%
PostTrain Bench
34.6%
71.9%
Program Bench
77.6%
40.0%
SWE Marathon
39.0%
69.2%
SWE-bench Pro
64.6%
31.55%
SpreadsheetBench 2
32.4%
84.6%
Terminal-Bench 2.1
88.8%
59.9%
Toolathlon
58.0%
76.2%
Toolathlon-Verified
74.9%
39.1%
WorldVQA
41.8%
17.0%
ZeroBench
17.0%