MMMU Pro
36 models tested · Updated 2026-06-10 · Verified sources only
Gemini 3.5 Flash leads at 83.6%
1
Google DeepMind · Blog/Google DeepMind · 2026-06-10
Highest among compared models. Up from Gemini 3 Flash (81.2).
83.6%
2
OpenAI · OpenAI Blog · 2026-04-23
With tools. 1pt over GPT-5.4 (82.1).
83.2%
3
OpenAI · Blog/OpenAI · 2026-07-09
Multimodal reasoning no-tools; edges GPT-5.5 (81.2%); 84.6% with tools.
83.0%
4
Moonshot AI · Blog/Moonshot AI · 2026-07-16
81.6%
5
OpenAI · Blog/OpenAI · 2026-03-05
Visual understanding and reasoning. Without tool use, reasoning effort xhigh.
81.2%
6
OpenAI · OpenAI Blog · 2026-04-23
No tools. Matches GPT-5.4, above Gemini 3.1 Pro (80.5).
81.2%
7
Anthropic · Blog/Anthropic · 2026-06-09
Score as cited in Kimi K3 blog (source: Anthropic official). Fable 5 hit fallbacks on 35% of tasks.
81.2%
8
Google · Blog/Google · 2025-12-17
Multimodal reasoning at Flash-tier cost. Competitive with Gemini 3 Pro.
81.0%
9
OpenAI · Blog/OpenAI · 2026-07-09
Terra nearly matches GPT-5.5 (81.2%) on MMMU Pro.
80.7%
10
Google · Official/Google DeepMind · 2026-02-19
Multimodal understanding without tools. Strong visual reasoning.
80.5%
11
Meta · Blog/Meta AI · 2026-04-08
Second-best MMMU Pro score, behind Gemini 3.1 Pro Preview (82.4%). Strong multimodal showing for Meta Superintelligence Labs debut model.
80.5%
12
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
Native multimodal model matching Gemini 3 Flash (81.2) on MMMU Pro.
79.4%
13
Alibaba · HuggingFace/Qwen · 2026-02-16
Native vision-language model. Strong multimodal reasoning.
79.0%
14
Anthropic · Blog/Anthropic · 2026-06-17
Score as cited in Kimi K3 blog (source: Anthropic official).
78.9%
15
Moonshot AI · Blog/Kimi · 2026-01-27
Strong multimodal reasoning for an open-weight model.
78.5%
16
OpenAI · Blog/OpenAI · 2026-07-09
Luna trails GPT-5.5 (81.2%) on MMMU Pro.
78.4%
17
Google · Model Card/Google · 2026-04-02
Multimodal vision benchmark. Up from 49.7% on Gemma 3 27B.
76.9%
18
Google · Model Card/Google · 2026-03-03
Budget-tier ($0.25/1M input) yet competitive on multimodal benchmarks.
76.8%
19
Alibaba · HuggingFace/Qwen · 2026-04-22
Strong multimodal reasoning for 27B.
75.8%
20
Alibaba · HuggingFace/Qwen · 2026-04-15
Matches Gemma 4 31B (76.9%). Strong vision-language reasoning for a MoE model.
75.3%
21
Anthropic · Blog/Google DeepMind · 2026-06-10
Below Gemini 3.5 Flash (83.6). Above Sonnet 4.6 (74.5).
75.2%
22
Alibaba · HuggingFace/Qwen · 2026-02-16
Multimodal understanding. Strong for a small model — competitive with some frontier scores from late 2025.
75.0%
23
Anthropic · Blog/Google DeepMind · 2026-06-10
Below all frontier models. Gemini 3.5 Flash leads (83.6).
74.5%
24
Google · Model Card/Google · 2026-04-02
MoE multimodal. 31B dense reaches 76.9%.
73.8%
25
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Vision reasoning; below Gemini 3.5 Flash (83.6) but strong for open-weights multimodal.
73.5%
26
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
Matches Inkling on MMMU Pro.
73.1%
27
Alibaba · HuggingFace/Qwen · 2026-03-02
9B params. Outperforms Gemini 2.5 Flash-Lite (59.7) on visual reasoning. Strong for its size class.
70.1%
28
Google DeepMind · HuggingFace/google-gemma-4-12b-it · 2026-05-23
Vision benchmark. Strong multimodal reasoning.
69.1%
29
PrismML · Whitepaper/PrismML · 2026-07-14
Ternary variant. Vision tower ships in 4-bit; multimodal agentic capability retained on-device.
68.96%
30
LG AI Research · HuggingFace/LGAI-EXAONE · 2026-04-14
Beats GPT-5 mini (67.3). Strong for an open-weight 33B VLM.
68.6%
31
PrismML · Whitepaper/PrismML · 2026-07-14
1-bit variant. Largest category drop (vision 72.6→59.6) — the 1-bit footprint trade-off lands heaviest on vision tasks.
60.48%
32
Vero Team · arxiv/2604.04917 · 2026-04-06
+3.9 over Qwen3-VL-8B-Instruct base. Open RL recipe with 600K samples across 59 datasets.
59.8%
33
Meta · HuggingFace/Meta · 2026-04-05
Official model card score. 17B active params, 128 experts.
59.6%
34
Google · Google Model Card · 2026-04-02
Multimodal understanding from a 4B model. Strong for edge deployment.
52.6%
35
Meta · HuggingFace/Meta · 2026-04-05
Official model card score. 17B active params, 16 experts, 10M context.
52.2%
36
Google · Model Card/Google · 2026-04-02
Multimodal reasoning at 2.3B active params. Natively multimodal from pretraining.
44.2%