MMLU Pro
55 models tested · Updated 2026-02-19 · Verified sources only
Gemini 3.1 Pro leads at 90.99%
1
Google · Blog/Google · 2026-02-19
Highest MMLU Pro score reported. Leads all models on this benchmark.
90.99%
2
OpenAI · Blog/OpenAI · 2025-08-05
Matches o4-mini on general knowledge. Runs on single 80GB GPU despite 116.8B params (5.1B active via MoE).
90.0%
3
Alibaba · Blog/Qwen · 2026-04-02
Leads MMLU-Pro leaderboard among available models as of early April 2026.
88.5%
4
Alibaba · HuggingFace/Qwen · 2026-02-16
Broad knowledge and reasoning. 397B total, 17B active params.
87.8%
5
ByteDance · Blog/ByteDance · 2026-03-10
Actually scores higher than Seed 2.0 Pro (87.0) on this benchmark.
87.7%
6
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning mode. Trails Gemini 3.1 Pro (91.0).
87.5%
7
Moonshot AI · HuggingFace/moonshotai · 2026-01-27
Strong knowledge reasoning from the 1T MoE open-source model. Thinking mode enabled.
87.1%
8
xAI · Artificial Analysis · 2025-07-09
Joint highest MMLU Pro at time of release. Verified by Artificial Analysis.
87.0%
9
ByteDance · Blog/ByteDance · 2026-02-14
Strong knowledge breadth. Seed 2.0 Lite variant scores slightly higher at 87.7.
87.0%
10
Anthropic · Blog/Anthropic · 2026-04-16
Up 5pp from Opus 4.6 (82.0%). Competitive with top proprietary models.
87.0%
11
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning, 13B activated. Near V4 Pro (87.5).
86.2%
12
Alibaba · HuggingFace/Qwen · 2026-04-22
Matches Gemma 4 31B despite being 4B smaller.
86.2%
13
Alibaba · Model Card/Qwen · 2026-02-24
Dense multimodal model, 262k native context.
86.1%
14
OpenAI · Blog/OpenAI · 2025-08-05
Only 3.6B active params. Matches o3-mini on general knowledge. Best-in-class for sub-30B models.
85.3%
15
Google · Model Card/Google · 2026-04-02
Updated from model card. Dense 31B instruction-tuned.
85.2%
16
Alibaba · HuggingFace/Qwen · 2026-04-16
Ties Gemma 4 31B on MMLU Pro with only 3B active parameters.
85.2%
17
DeepSeek · Blog/DeepSeek · 2025-12-01
Strong knowledge benchmark. Competitive with top models.
85.0%
18
Xiaomi · arxiv/2601.02780 · 2026-01-06
15B active MoE. Pre-trained on 27T tokens with novel distillation approach.
84.9%
19
StepFun · HuggingFace/stepfun-ai · 2026-02-02
Open-weight MoE with 11B active params. Strong general knowledge for its size.
84.4%
20
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
84.4%
21
NVIDIA · Blog/NVIDIA · 2026-03-11
Lags Qwen3.5 122B (86.70%) on MMLU-Pro but leads on throughput efficiency.
83.73%
22
Solid general reasoning for a 13B-active MoE. Competitive with much larger models.
83.4%
23
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
83.3%
24
LG AI Research · HuggingFace/LGAI-EXAONE · 2026-04-14
Matches GPT-5 mini (83.3). First open-weight VLM to match frontier-class reasoning benchmarks.
83.3%
25
OpenAI · arxiv/2604.08644 · 2026-04-14
From EXAONE 4.5 technical report. Tied with EXAONE 4.5 33B.
83.3%
26
Google · Model Card/Google · 2026-04-02
MoE with 3.8B active params. Strong efficiency.
82.6%
27
Alibaba · HuggingFace/Qwen · 2026-02-28
9B params beating GPT-OSS-120B (80.8). Best-in-class for small open models.
82.5%
28
Anthropic · Blog/Anthropic · 2026-02-05
Strong general knowledge, but trails Gemini 3.1 Pro and GPT-5 on this benchmark.
82.0%
29
Sarvam AI · HuggingFace/sarvamai · 2026-03-06
India's first domestically-trained 105B model. MoE with 10.3B active params. Apache 2.0.
81.7%
30
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
81.7%
31
JD.com · Paper/JD.com (arXiv) · 2026-04-03
48B MoE, 2.7B active params. Competitive with 9B-class models at fraction of compute.
81.6%
32
Qwen · arxiv/2603.00729 · 2026-02-28
Nearly matches Qwen3-Next (80.89) on MMLU Pro despite being coding-specialized with 3B active params.
80.52%
33
Meta · HuggingFace/Meta · 2026-04-05
17B active params, 128 experts, 400B total. Instruction-tuned score from official model card.
80.5%
34
NVIDIA · arxiv/2603.19220 · 2026-03-19
Lower MMLU-Pro than Qwen3.5-35B-A3B (85.3) but stronger on math/coding, showing reasoning-knowledge tradeoff.
79.8%
35
Anthropic · Blog/Anthropic · 2026-02-17
Broad knowledge and reasoning. Anthropic official benchmarks page.
79.2%
36
Alibaba · HuggingFace/Alibaba · 2026-03-02
Strong general knowledge for a 4B open-weight model.
79.1%
37
NVIDIA · Blog/NVIDIA · 2025-12-15
Hybrid MoE model, 30B params with 3B active. Mamba-Transformer architecture with 1M token context. Open-weight under NVIDIA license.
78.3%
38
Mistral · Blog/Mistral · 2026-03-16
Efficient small model. Competitive with much larger models using less compute.
78.0%
39
ServiceNow · arxiv/2604.02007 · 2026-04-02
15B open-weight. Only 1.9K mean output tokens, roughly half of next most efficient model.
77.3%
40
Google DeepMind · HuggingFace/google-gemma-4-12b-it · 2026-05-23
12B dense unified multimodal model. Competitive with much larger Gemma 3 27B (67.6).
77.2%
41
Microsoft · HuggingFace/Microsoft · 2025-04-30
SFT+RL variant. 14B params, MIT license, 32k context.
76.0%
42
Independent · arxiv/2604.08477 · 2026-04-09
RLVR fine-tuned Qwen3-4B. Matches Qwen3-8B MMLU Pro despite half the params.
76.0%
43
Meta · HuggingFace/Meta · 2026-04-05
Instruction-tuned score from official model card.
74.3%
44
Microsoft · HuggingFace/Microsoft · 2025-04-30
14B dense model, MIT license. Strong for its size class.
74.3%
45
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
69.8%
46
Google · Google Model Card · 2026-04-02
Official model card score. Up from 60.0% (tweet source).
69.4%
47
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
67.1%
48
Tencent · GitHub/Tencent-Hunyuan-Hy3-preview · 2026-04-23
Hy3 Preview base model (pre-instruct). 21B active MoE matching Kimi-K2 base.
65.76%
49
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
63.4%
50
arxiv · arxiv/2604.08477 · 2026-04-09
RLVR framework that adapts instruction-tuning data for RL. SuperNova-4B outperforms Qwen3-4B by 8.8pp average on MMLU-Pro, BBH, Zebralogic, MATH500. 100+ controlled RL experiments. Also evaluates on B
61.5%
51
Google · Model Card/Google · 2026-04-02
Ultra-compact 2.3B active params (5.1B total) with PLE. Fits in 1.5GB quantized. Strong for its size class.
60.0%
52
arxiv · arxiv/2604.08477 · 2026-04-09
RLVR framework that adapts instruction-tuning data for RL. SuperNova-4B outperforms Qwen3-4B by 8.8pp average on MMLU-Pro, BBH, Zebralogic, MATH500. 100+ controlled RL experiments. Also evaluates on B
56.2%
53
arxiv · arxiv/2604.07023 · 2026-04-09
Lightweight fine-tuning that teaches AR models to predict multiple tokens per forward pass. At 7B on Qwen2.5, MARS improves GSM8K by 4.5 points and HumanEval by 3.0 points over AR SFT while enabling 1
44.4%
54
arxiv · arxiv/2603.13398 · 2026-03-11
4B end-to-end OCR model that ranks #1 on OmniDocBench among end-to-end models. Introduces "Layout-as-Thought" for structured layout representations. Outperforms Qwen3-VL-4B on ChartQA (+4.8) and Chart
33.8%
55
arxiv · arxiv/2604.07023 · 2026-04-09
Lightweight fine-tuning that teaches AR models to predict multiple tokens per forward pass. At 7B on Qwen2.5, MARS improves GSM8K by 4.5 points and HumanEval by 3.0 points over AR SFT while enabling 1
12.4%