benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
arxiv
JoyAI-LLM Flash
7 benchmarks
HumanEval
#4 of 19
94.5%
HumanEval
#4 of 19
87.5%
MMLU Pro
#31 of 55
81.6%
GPQA Diamond
#77 of 118
74.5%
AIME
#69 of 117
72.9%
LiveCodeBench
#37 of 53
65.6%
SWE-bench Verified
#59 of 86
62.6%