benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
Anthropic
Claude Opus 4.8
36 benchmarks
HMMT Feb 2026
#2 of 5
96.7%
HMMT Nov 2025
#1 of 4
96.5%
AIME
#19 of 117
95.7%
GPQA Diamond
#10 of 118
93.6%
GPQA Diamond
#10 of 118
91.0%
OmniDocBench
#27 of 31
87.9%
MATH-Vision
#5 of 15
86.7%
Terminal-Bench 2.1
#7 of 27
85.0%
Terminal-Bench 2.1
#7 of 27
84.6%
BrowseComp
#15 of 45
84.3%
MCP-Atlas
#5 of 17
83.6%
BabyVision
#5 of 5
81.2%
CharXiv
#13 of 23
80.5%
MMMU Pro
#14 of 36
78.9%
MCP-Atlas
#5 of 17
77.8%
Toolathlon-Verified
#2 of 7
76.2%
FrontierSWE
#3 of 8
75.1%
Program Bench
#4 of 6
71.9%
SWE-bench Pro
#4 of 43
69.2%
FrontierSWE
#3 of 8
66.7%
OfficeQA Pro
#2 of 6
63.9%
Toolathlon
#2 of 7
59.9%
DeepSWE
#11 of 22
59.0%
Humanity's Last Exam
#6 of 72
57.9%
OSWorld 2.0
#2 of 10
54.8%
Humanity's Last Exam
#6 of 72
49.8%
JobBench
#3 of 6
48.4%
PerceptionBench
#5 of 5
47.2%
MLS Bench
#4 of 6
42.8%
SWE Marathon
#2 of 6
40.0%
APEX-Agents
#4 of 7
39.4%
WorldVQA
#4 of 5
39.1%
PostTrain Bench
#5 of 6
34.1%
SpreadsheetBench 2
#4 of 6
31.55%
AutomationBench
#4 of 6
27.2%
ZeroBench
#5 of 5
17.0%