benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
OpenAI
GPT-5
11 benchmarks
AIME
#24 of 117
94.6%
DocVQA
#13 of 15
91.5%
AIME
#24 of 117
90.0%
Aider Polyglot
#1 of 12
88.0%
HumanEval
#8 of 19
85.0%
MMMU
#4 of 27
84.2%
Aider Polyglot
#1 of 12
81.3%
InfoVQA
#4 of 8
79.0%
SWE-bench Verified
#31 of 86
74.9%
SWE-bench Verified
#31 of 86
66.0%
Humanity's Last Exam
#63 of 72
20.0%