benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
MLE-Bench leaderboard
MLE-Bench
6 models tested · Updated 2026-07-21 · Verified sources only
Claude Sonnet 5
leads at
66.9%
1
Claude Sonnet 5
Anthropic ·
Blog/Google DeepMind (cross-ref)
· 2026-07-21
Google's measurement. Highest in the comparison table.
66.9%
2
Gemini 3.6 Flash
Google DeepMind ·
Blog/Google DeepMind
· 2026-07-21
Beats GPT-5.6 (47.6) and Grok 4.5 (43.2) on ML engineering; near Sonnet 5.
63.9%
3
Gemini 3.5 Flash
Google DeepMind ·
Blog/Google DeepMind (cross-ref)
· 2026-07-21
Google official. ML engineering benchmark.
49.7%
4
GPT-5.6 Luna
OpenAI ·
Blog/Google DeepMind (cross-ref)
· 2026-07-21
Google's measurement. Surprisingly low vs Gemini 3.6 Flash (63.9).
47.6%
5
Grok 4.5
xAI ·
Blog/Google DeepMind
· 2026-07-21
From Google DeepMind eval table. Behind Gemini 3.6 Flash (63.9%) and Claude Sonnet 5 (66.9%).
43.2%
6
Gemini 3.1 Pro
Google DeepMind ·
Blog/Google DeepMind (cross-ref)
· 2026-07-21
Google official. ML engineering.
42.6%