benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
Anthropic
Claude Fable 5
30 benchmarks
MATH-Vision
#2 of 15
94.8%
GPQA Diamond
#17 of 118
92.6%
BabyVision
#1 of 5
90.5%
OmniDocBench
#19 of 31
89.8%
CharXiv
#2 of 23
88.9%
BrowseComp
#8 of 45
88.0%
Terminal-Bench 2.1
#4 of 27
88.0%
FrontierSWE
#1 of 8
86.6%
OSWorld-Verified
#1 of 6
85.0%
MCP-Atlas
#1 of 17
84.7%
Terminal-Bench 2.1
#4 of 27
84.6%
MMMU Pro
#7 of 36
81.2%
SWE-bench Pro
#1 of 43
80.3%
Toolathlon-Verified
#1 of 7
77.9%
Program Bench
#3 of 6
76.8%
DeepSWE
#3 of 22
70.0%
OfficeQA Pro
#1 of 6
69.9%
Humanity's Last Exam
#2 of 72
59.0%
JobBench
#1 of 6
57.4%
PerceptionBench
#3 of 5
57.2%
WorldVQA
#1 of 5
56.7%
Humanity's Last Exam
#2 of 72
53.3%
MLS Bench
#1 of 6
49.9%
APEX-Agents
#1 of 7
43.3%
PostTrain Bench
#1 of 6
41.4%
SWE Marathon
#4 of 6
35.0%
SpreadsheetBench 2
#2 of 6
34.7%
FrontierCode Diamond
#1 of 1
29.3%
AutomationBench
#3 of 6
29.1%
ZeroBench
#2 of 5
23.0%