MLE-Bench
6 models tested · Updated 2026-07-21 · Verified sources only
Claude Sonnet 5 leads at 66.9%
1
Anthropic · Blog/Google DeepMind (cross-ref) · 2026-07-21
Google's measurement. Highest in the comparison table.
66.9%
2
Google DeepMind · Blog/Google DeepMind · 2026-07-21
Beats GPT-5.6 (47.6) and Grok 4.5 (43.2) on ML engineering; near Sonnet 5.
63.9%
3
Google DeepMind · Blog/Google DeepMind (cross-ref) · 2026-07-21
Google official. ML engineering benchmark.
49.7%
4
OpenAI · Blog/Google DeepMind (cross-ref) · 2026-07-21
Google's measurement. Surprisingly low vs Gemini 3.6 Flash (63.9).
47.6%
5
xAI · Blog/Google DeepMind · 2026-07-21
From Google DeepMind eval table. Behind Gemini 3.6 Flash (63.9%) and Claude Sonnet 5 (66.9%).
43.2%
6
Google DeepMind · Blog/Google DeepMind (cross-ref) · 2026-07-21
Google official. ML engineering.
42.6%