benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
Moonshot AI
Kimi K3
28 benchmarks
MATH-Vision
#3 of 15
94.3%
GPQA Diamond
#11 of 118
93.5%
BrowseComp
#2 of 45
91.2%
OmniDocBench
#11 of 31
91.1%
BrowseComp
#2 of 45
90.4%
Terminal-Bench 2.1
#3 of 27
88.3%
BabyVision
#3 of 5
85.7%
CharXiv
#5 of 23
84.8%
MCP-Atlas
#2 of 17
84.2%
MMMU Pro
#4 of 36
81.6%
FrontierSWE
#2 of 8
81.2%
Program Bench
#1 of 6
77.8%
Toolathlon-Verified
#5 of 7
73.2%
DeepSWE
#5 of 22
69.0%
DeepSWE
#5 of 22
67.5%
DeepSWE
#5 of 22
67.3%
OfficeQA Pro
#3 of 6
63.3%
PerceptionBench
#2 of 5
58.5%
JobBench
#2 of 6
52.9%
WorldVQA
#2 of 5
51.0%
MLS Bench
#2 of 6
48.3%
Humanity's Last Exam
#31 of 72
43.5%
SWE Marathon
#1 of 6
42.0%
APEX-Agents
#2 of 7
41.0%
PostTrain Bench
#2 of 6
36.6%
SpreadsheetBench 2
#1 of 6
34.8%
AutomationBench
#1 of 6
30.8%
ZeroBench
#1 of 5
23.0%