MMMLU
7 models tested · Updated 2026-04-16 · Verified sources only
Claude Opus 4.7 leads at 92.0%
1
Anthropic · Blog/Anthropic · 2026-04-16
Up from Opus 4.6 (91.1%). Multilingual knowledge benchmark.
92.0%
2
Anthropic · arxiv/Mythos-System-Card · 2026-04-07
Multilingual MMLU. Slightly below Gemini 3.1 Pro (92.6-93.6%).
91.1%
3
Google · HuggingFace/Google DeepMind · 2026-04-02
Massive multilingual MMLU. 30.7B dense.
88.4%
4
Google · HuggingFace/Google DeepMind · 2026-04-02
MoE 25.2B total, 3.8B active.
86.3%
5
Google DeepMind · HuggingFace/google-gemma-4-12b-it · 2026-05-23
Multilingual MMLU.
83.4%
6
Google · HuggingFace/Google DeepMind · 2026-04-02
4.5B effective params.
76.6%
7
Google · HuggingFace/Google DeepMind · 2026-04-02
2.3B effective params.
67.4%