GPT-5.5 vs Claude Fable 5
32 shared benchmarks
2
wins
1
ties
29
wins
38.5%
APEX-Agents
43.3%
46.9%
Agents' Last Exam
40.5%
22.7%
AutomationBench
29.1%
83.6%
BabyVision
90.5%
84.4%
BrowseComp
88.0%
84.1%
CharXiv
88.9%
67.0%
DeepSWE
70.0%
47.9%
ExploitBench
78.0%
51.7%
FrontierMath
87.0%
35.4%
FrontierMath Tier 4
87.8%
64.9%
FrontierSWE
86.6%
93.5%
GPQA Diamond
92.6%
41.4%
Humanity's Last Exam
53.3%
38.3%
JobBench
57.4%
92.2%
MATH-Vision
94.8%
82.8%
MCP-Atlas
84.7%
35.5%
MLS Bench
49.9%
81.2%
MMMU Pro
81.2%
78.7%
OSWorld-Verified
85.0%
60.9%
OfficeQA Pro
69.9%
89.4%
OmniDocBench
89.8%
55.8%
PerceptionBench
57.2%
28.4%
PostTrain Bench
41.4%
70.8%
Program Bench
76.8%
14.0%
SWE Marathon
35.0%
58.6%
SWE-bench Pro
80.3%
29.05%
SpreadsheetBench 2
34.7%
83.4%
Terminal-Bench 2.1
84.6%
55.6%
Toolathlon
61.7%
73.5%
Toolathlon-Verified
77.9%
38.5%
WorldVQA
56.7%
22.0%
ZeroBench
23.0%