benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
OpenAI
GPT-5.6 Sol
35 benchmarks
MATH-Vision
#1 of 15
95.8%
GPQA Diamond
#2 of 118
94.6%
GPQA Diamond
#2 of 118
94.1%
BrowseComp
#3 of 45
90.4%
FrontierMath
#1 of 11
89.0%
BabyVision
#2 of 5
88.9%
Terminal-Bench 2.1
#2 of 27
88.8%
OmniDocBench
#29 of 31
85.8%
CharXiv
#6 of 23
84.6%
MCP-Atlas
#4 of 17
83.6%
FrontierMath
#1 of 11
83.0%
MMMU Pro
#3 of 36
83.0%
Program Bench
#2 of 6
77.6%
Toolathlon-Verified
#3 of 7
74.9%
ExploitBench
#2 of 2
73.5%
DeepSWE
#1 of 22
73.0%
DeepSWE
#1 of 22
72.7%
FrontierSWE
#5 of 8
71.3%
SEC-Bench Pro
#2 of 2
71.2%
SWE-bench Pro
#6 of 43
64.6%
OfficeQA Pro
#4 of 6
63.2%
OSWorld 2.0
#1 of 10
62.6%
PerceptionBench
#1 of 5
59.7%
JobBench
#4 of 6
46.5%
MLS Bench
#3 of 6
46.2%
Humanity's Last Exam
#28 of 72
44.5%
WorldVQA
#3 of 5
41.8%
APEX-Agents
#3 of 7
39.9%
SWE Marathon
#3 of 6
39.0%
PostTrain Bench
#3 of 6
34.6%
ExploitGym
#1 of 1
33.7%
SpreadsheetBench 2
#3 of 6
32.4%
AutomationBench
#2 of 6
29.7%
ZeroBench
#4 of 5
17.0%
ARC-AGI 3
#1 of 8
7.78%