GPT-5.6 Sol vs Claude Fable 5
30 shared benchmarks
11
wins
0
ties
19
wins
39.9%
APEX-Agents
43.3%
52.7%
Agents' Last Exam
40.5%
29.7%
AutomationBench
29.1%
88.9%
BabyVision
90.5%
90.4%
BrowseComp
88.0%
84.6%
CharXiv
88.9%
72.7%
DeepSWE
70.0%
73.5%
ExploitBench
78.0%
83.0%
FrontierMath
87.0%
71.3%
FrontierSWE
86.6%
94.6%
GPQA Diamond
92.6%
44.5%
Humanity's Last Exam
53.3%
46.5%
JobBench
57.4%
95.8%
MATH-Vision
94.8%
83.6%
MCP-Atlas
84.7%
46.2%
MLS Bench
49.9%
83.0%
MMMU Pro
81.2%
63.2%
OfficeQA Pro
69.9%
85.8%
OmniDocBench
89.8%
59.7%
PerceptionBench
57.2%
34.6%
PostTrain Bench
41.4%
77.6%
Program Bench
76.8%
39.0%
SWE Marathon
35.0%
64.6%
SWE-bench Pro
80.3%
32.4%
SpreadsheetBench 2
34.7%
88.8%
Terminal-Bench 2.1
84.6%
58.0%
Toolathlon
61.7%
74.9%
Toolathlon-Verified
77.9%
41.8%
WorldVQA
56.7%
17.0%
ZeroBench
23.0%