benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
MMMLU leaderboard
MMMLU
7 models tested · Updated 2026-04-16 · Verified sources only
Claude Opus 4.7
leads at
92.0%
1
Claude Opus 4.7
Anthropic ·
Blog/Anthropic
· 2026-04-16
Up from Opus 4.6 (91.1%). Multilingual knowledge benchmark.
92.0%
2
Claude Opus 4.6
Anthropic ·
arxiv/Mythos-System-Card
· 2026-04-07
Multilingual MMLU. Slightly below Gemini 3.1 Pro (92.6-93.6%).
91.1%
3
Gemma 4 31B
Google ·
HuggingFace/Google DeepMind
· 2026-04-02
Massive multilingual MMLU. 30.7B dense.
88.4%
4
Gemma 4 26B A4B
Google ·
HuggingFace/Google DeepMind
· 2026-04-02
MoE 25.2B total, 3.8B active.
86.3%
5
Gemma 4 12B
Google DeepMind ·
HuggingFace/google-gemma-4-12b-it
· 2026-05-23
Multilingual MMLU.
83.4%
6
Gemma 4 E4B
Google ·
HuggingFace/Google DeepMind
· 2026-04-02
4.5B effective params.
76.6%
7
Gemma 4 E2B
Google ·
HuggingFace/Google DeepMind
· 2026-04-02
2.3B effective params.
67.4%