GPQA Diamond
118 models tested · Updated 2026-04-07 · Verified sources only
Claude Mythos Preview leads at 94.6%
1
Anthropic · Blog/Anthropic · 2026-04-07
New SOTA. Tied with Gemini 3.1 Pro (94.3%). PhD-level science reasoning.
94.6%
2
OpenAI · Blog/OpenAI · 2026-07-09
Ties Claude Mythos Preview at the top of GPQA Diamond; edges GPT-5.5 (93.6%).
94.6%
3
OpenAI · OpenAI Blog · 2026-04-23
Wait, this is GPT-5.4 Pro score from the table. Skip.
94.4%
4
OpenAI · OpenAI Blog · 2026-04-23
Near top of leaderboard, above GPT-5.5 (93.6).
94.4%
5
Google · PCMag/Google DeepMind · 2026-02-19
Highest GPQA Diamond score ever recorded. Leads 13 of 16 major benchmarks per Google DeepMind.
94.3%
6
Anthropic · Blog/Anthropic · 2026-04-16
Surpasses Gemini 3.1 Pro (94.3%) and GPT-5.4 (92.0%) on GPQA Diamond per multiple sources.
94.2%
7
Anthropic · Blog/OpenAI · 2026-06-09
Slightly below Mythos Preview (94.6) and GPT-5.6 Sol (94.6); above Fable 5 (92.6) — no safeguard fallback on reasoning.
94.1%
8
OpenAI · Blog/OpenAI · 2026-07-09
Score as cited in Kimi K3 blog (source: OpenAI official).
94.1%
9
OpenAI · OpenAI Blog · 2026-04-23
Slightly below GPT-5.4 Pro (94.4) and Gemini 3.1 Pro (94.3).
93.6%
10
Anthropic · Blog/Z.ai · 2026-06-13
Ties GPT-5.5. Below Gemini 3.1 Pro (94.3).
93.6%
11
Moonshot AI · Blog/Moonshot AI · 2026-07-16
93.5%
12
OpenAI · Blog/OpenAI · 2026-06-26
Score as cited in Kimi K3 blog (source: OpenAI official).
93.5%
13
OpenAI · Blog/OpenAI · 2025-12-11
OpenAI Pro-tier model. Surpasses PhD experts (69.7%) by large margin. Just below Gemini 3.1 Pro (94.3%).
93.2%
14
MiniMax · Blog/Z.ai · 2026-06-13
Strong reasoning. Ties deep V4 Pro tier.
93.0%
15
OpenAI · Blog/OpenAI · 2026-07-09
Terra trails Sol (94.6%) but beats Opus 4.8 (92%).
92.9%
16
OpenAI · Blog/OpenAI · 2026-03-05
GPT-5.4 scores 92.8% on GPQA Diamond, marginal improvement over GPT-5.2 at 92.4%
92.8%
17
Anthropic · Blog/Anthropic · 2026-06-09
Score as cited in Kimi K3 blog (source: Anthropic official). Fable 5 hit fallbacks on 35% of tasks.
92.6%
18
GPT-5.2 Thinking variant. +4.3% over GPT-5.1. No tools, max reasoning effort.
92.4%
19
OpenAI · Blog/OpenAI · 2026-07-09
Luna beats Opus 4.8 (92%) on GPQA Diamond at lowest tier.
92.3%
20
Google · Blog/Google · 2025-11-18
Announced with Gemini 3 launch. PhD-level reasoning, top of LMArena at 1501 Elo.
91.9%
21
OpenAI · Artificial Analysis · 2026-02-05
Third-highest GPQA Diamond score. Codex-native agent pairing frontier coding with general reasoning.
91.5%
22
Graduate-level science reasoning. Within 1.1 points of GPT-5.2.
91.3%
23
Z.ai · Blog/Z.ai · 2026-06-13
Up from GLM-5.1 (86.2). Competitive with closed frontier (GPT-5.5 93.6).
91.2%
24
Anthropic · Blog/Anthropic · 2026-06-17
Score as cited in Kimi K3 blog (source: Anthropic official).
91.0%
25
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
Strong showing from open-weight MoE. Up from K2.5 (87.6%) but below Opus 4.7 (94.2%).
90.5%
26
Alibaba · Blog/Qwen · 2026-04-02
Qwen 3.6 Plus leads on GPQA with 90.4%
90.4%
27
Google · Blog/Google · 2025-12-17
Frontier PhD-level reasoning at Flash-tier pricing and latency. Matches larger models.
90.4%
28
Tencent · HuggingFace/tencent-Hy3-modelcard · 2026-07-06
GPQA Diamond. Rivals flagship open and closed models at 21B active params.
90.4%
29
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning mode. Best open-source model, trails closed-source frontier by ~4pt.
90.1%
30
Alibaba · Blog/Z.ai · 2026-06-13
Below GLM-5.2 (91.2). Mid-pack.
90.0%
31
Anthropic · Blog/Anthropic · 2026-02-17
Averaged over 10 trials with adaptive thinking. Within 2.5 pts of GPT-5.2.
89.9%
32
Meta · Blog/Meta · 2026-04-08
Strong PhD-level reasoning. Ahead of Gemini 3.1 Pro (90.8%) and GLM-5.1 (86.2%).
89.5%
33
ByteDance · Blog/ByteDance · 2026-02-14
ByteDance frontier model. Gold medals on ICPC, IMO, CMO. Competes with GPT-5.2.
88.9%
34
Alibaba · HuggingFace/Qwen · 2026-02-16
397B total, 17B active MoE. Native multimodal with vision. Top open-weight model on reasoning.
88.4%
35
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
Beats the larger Inkling (87.2) on GPQA Diamond.
88.3%
36
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning, 13B activated params. Impressive for small model.
88.1%
37
xAI · Artificial Analysis · 2025-07-09
All-time high GPQA Diamond at time of release. Verified by Artificial Analysis.
88.0%
38
OpenAI · Blog/OpenAI · 2026-03-17
Small/cheap GPT-5.4 variant approaching full GPT-5.4 (92.8%) on reasoning. 2x faster than GPT-5 mini.
88.0%
39
Alibaba · HuggingFace/Qwen · 2026-04-22
Top-tier reasoning for 27B, beating Gemma 4 31B.
87.8%
40
OpenAI · Blog/OpenAI · 2025-04-16
OpenAI reasoning model. Strong scientific knowledge but below Gemini 3.1 Pro (94.3%) and Claude Opus 4.6 (91.3%).
87.7%
41
Moonshot AI · HuggingFace/moonshotai · 2026-01-27
1T MoE, 32B active. Open-source. Averaged over 8 runs. Strong but trails Gemini 3.1 Pro (94.3) and Claude Opus 4.6 (91.3).
87.6%
42
xAI · Blog/XAI · 2025-07-09
Official xAI benchmark (no tools). Slightly below vals.ai independent test of 88.0%.
87.5%
43
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Below frontier (GPT-5.6 at 94.1) but strong for open-weights. Controllable thinking effort at 0.99.
87.2%
44
Google · Model Card/Google · 2026-03-03
Lightweight model rivaling larger Gemini 3 series on reasoning. 2.5x faster than 2.5 Flash.
86.9%
45
MiniMax · Artificial Analysis · 2026-03-18
Self-improving model that ran 100+ autonomous optimization cycles during training. Competitive with Kimi K2.5 on science reasoning.
86.2%
46
Zhipu AI · Blog/Z.AI · 2026-04-07
Strong reasoning for a 744B MoE open-weight model. Behind Gemini 3.1 Pro (94.3) and GPT-5.4 (92.8).
86.2%
47
Zhipu AI · Paper/Zhipu AI (arxiv:2602.15763) · 2026-02-11
Open-weight 744B MoE (40B active). Competitive with proprietary models, trails only Gemini 3 Pro and GPT-5.2.
86.0%
48
Alibaba · HuggingFace/Qwen · 2026-04-16
Open-weight MoE with 3B active params matching much larger models on graduate-level reasoning.
86.0%
49
Alibaba · X/@rohanpaul_ai · 2026-02-24
Highest recorded result among open-weights models under 40B parameters on GPQA Diamond
85.8%
50
Google · X/@rohanpaul_ai · 2026-04-02
2nd-highest among open-weights models under 40B parameters on GPQA Diamond
85.7%
51
DeepSeek · Blog/DeepSeek · 2025-12-01
Deep reasoning variant. Par with Gemini-3.0-Pro on scientific reasoning.
85.7%
52
Zhipu AI · NVIDIA/Zhipu Official · 2025-12-22
Open-weight, 400B params. Strong science reasoning for an open model.
85.7%
53
Alibaba · HuggingFace/Alibaba · 2026-02-20
Dense 27B model, natively multimodal. Competitive with much larger frontier models on graduate-level reasoning.
85.5%
54
MiniMax · Blog/MiniMax · 2026-02-12
Strong graduate-level reasoning. Between Gemini 3 Pro (91.0) and Claude Sonnet 4.5 (83.0).
85.2%
55
ByteDance · Blog/ByteDance · 2026-03-10
Strong for a mid-tier model. 262K context window, multimodal.
85.1%
56
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
85.0%
57
Google · Model Card/Google · 2026-04-02
From Google model card. Dense 31B instruction-tuned.
84.3%
58
Xiaomi · arxiv/2601.02780 · 2026-01-06
15B active MoE. Post-training score, surpasses many much larger open models.
84.3%
59
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
84.3%
60
StepFun · HuggingFace/stepfun-ai · 2026-02-02
Open-weight MoE (196B total, 11B active). From official model card.
83.5%
61
OpenAI · Blog/OpenAI · 2026-03-17
Smallest GPT-5.4 variant. Strong GPQA for its cost tier ($0.20/1M input tokens).
82.8%
62
DeepSeek · Blog/DeepSeek · 2025-12-01
Graduate-level science reasoning. Solid mid-frontier performance.
82.4%
63
Google · Google — Gemma 4 Blog · 2026-04-02
MoE with 3.8B active params. Outperforms Sonnet 4.6 (74%). Runs on 24GB hardware.
82.3%
64
Alibaba · HuggingFace/Qwen · 2026-03-02
9B params. Outperforms GPT-OSS-120B (80.1). Best reasoning-per-parameter ratio among open models.
81.7%
65
OpenAI · Blog/OpenAI · 2025-04-16
Strong scientific reasoning from OpenAI cost-efficient reasoning model. Competitive with much larger models.
81.4%
66
LG AI Research · HuggingFace/LGAI-EXAONE · 2026-04-14
Open-weight VLM with strong reasoning. Below Qwen3.5 27B (85.5) but competitive with much larger MoE models.
80.5%
67
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
79.6%
68
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
79.5%
69
NVIDIA · Blog/NVIDIA · 2026-03-11
Open-weight MoE, 120B total/12B active params. Without tools. 82.70% with tools. NVFP4 quantization.
79.42%
70
Google DeepMind · HuggingFace/google-gemma-4-12b-it · 2026-05-23
GPQA Diamond. Near GPT-4-class reasoning at 12B params.
78.8%
71
Sarvam AI · HuggingFace/sarvamai · 2026-03-06
India's first domestically-trained 105B model. MoE with 10.3B active params. Apache 2.0.
78.7%
72
Trinity-Large-Thinking: 398B MoE, only 13B active params. Open-weight, Apache 2.0.
76.3%
73
Alibaba · HuggingFace/Alibaba · 2026-03-02
Remarkable for a 4B model. Matches larger Qwen3-80B-A3B on reasoning.
76.2%
74
NVIDIA · arxiv/2603.19220 · 2026-03-19
Competitive GPQA Diamond for 3B active params; slightly below Qwen3.5-35B-A3B (84.2) but with 10x fewer active parameters.
76.1%
75
Zhipu AI · HuggingFace/Zhipu · 2026-01-15
30B-A3B MoE, 3.6B active params. Competitive with much larger models on graduate-level reasoning.
75.2%
76
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
75.0%
77
arxiv · arxiv/2604.03044 · 2026-04-03
New 48B MoE model (2.7B active) trained on 20T tokens with novel FiberPO RL algorithm. Open-sourced on HuggingFace. Competitive with much larger models on SWE-bench Verified (62.6%), LiveCodeBench v6
74.5%
78
Qwen · arxiv/2603.00729 · 2026-02-28
Strong GPQA Diamond for a coding-specialized model with 3B active params.
74.49%
79
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
74.0%
80
OpenAI · HuggingFace/openai/gpt-oss-120b · 2025-08-05
Primary source (HuggingFace model card, arxiv paper). Scores up to 80.9% reported with tool augmentation.
73.5%
81
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
72.5%
82
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
72.4%
83
Mistral · Blog/Mistral · 2026-03-16
71.2% on GPQA Diamond with efficient output length. Released March 16, 2026.
71.2%
84
Meta · HuggingFace/Meta · 2026-04-05
17B active params, 128 experts, 400B total. Released alongside Scout on April 5. MoE architecture, natively multimodal.
69.8%
85
ServiceNow · arxiv/2604.02007 · 2026-04-02
15B open-weight model. Highest accuracy among comparable-size models while using roughly half the tokens.
69.8%
86
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
69.5%
87
Microsoft · HuggingFace/Microsoft · 2025-04-30
14B SFT+RL model. Outperforms DeepSeek-R1-Distill-70B on most benchmarks.
68.9%
88
Alibaba Cloud · arxiv/2601.09088 · 2026-01-15
4B model achieving near-70B-class reasoning. Novel distillation approach with Temperature-scheduled Learning.
68.4%
89
arxiv · arxiv/2604.08299 · 2026-04-09
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
67.17%
90
OpenAI · HuggingFace/openai/gpt-oss-20b · 2025-08-05
Without tool augmentation. HuggingFace model card primary source. 20.9B total params, 3.6B active.
67.1%
91
Microsoft · HuggingFace/Microsoft · 2025-04-30
14B model beating o1-mini (60.0%) and QwQ-32B (59.5%) on graduate-level science.
65.8%
92
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
65.6%
93
arxiv · arxiv/2604.05424 · 2026-04-06
MCTS with metacognitive reflection improves reasoning. On Qwen3-30B-A3B, PRISM-MCTS achieves 65.15% on GPQA Diamond and 93.28% on MATH500.
65.15%
94
arxiv · arxiv/2604.05424 · 2026-04-06
MCTS with metacognitive reflection improves reasoning. On Qwen3-30B-A3B, PRISM-MCTS achieves 65.15% on GPQA Diamond and 93.28% on MATH500.
65.08%
95
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
64.4%
96
arxiv · arxiv/2604.08299 · 2026-04-09
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
61.62%
97
Google · Google Model Card · 2026-04-02
Strong graduate-level reasoning for a 4B edge model.
58.6%
98
Meta · HuggingFace/Meta · 2026-04-05
17B active params, 16 experts. 10M context window. Smaller expert count than Maverick.
57.2%
99
arxiv · arxiv/2604.08468 · 2026-04-09
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
56.1%
100
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
52.5%
101
Qwen · arxiv/2604.05355 · 2026-04-07
GPQA Diamond baseline for 8B model; ETR reaches 57.1, showing efficient reasoning gains.
52.0%
102
arxiv · arxiv/2604.08468 · 2026-04-09
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
51.5%
103
arxiv · arxiv/2604.08468 · 2026-04-09
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
48.9%
104
arxiv · arxiv/2604.05424 · 2026-04-06
MCTS with metacognitive reflection improves reasoning. On Qwen3-30B-A3B, PRISM-MCTS achieves 65.15% on GPQA Diamond and 93.28% on MATH500.
47.55%
105
arxiv · arxiv/2604.05424 · 2026-04-06
MCTS with metacognitive reflection improves reasoning. On Qwen3-30B-A3B, PRISM-MCTS achieves 65.15% on GPQA Diamond and 93.28% on MATH500.
45.8%
106
Qwen · arxiv/2604.05355 · 2026-04-07
GPQA Diamond baseline for 4B model; ETR reaches 53.5 showing room for improvement with better reasoning efficiency.
43.9%
107
Google · Model Card/Google · 2026-04-02
Ultra-compact model with Per-Layer Embeddings. Competitive with models 10x its size on science reasoning.
43.4%
108
arxiv · arxiv/2604.08299 · 2026-04-09
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
40.91%
109
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
38.4%
110
arxiv · arxiv/2604.08299 · 2026-04-09
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
35.35%
111
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
34.9%
112
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
29.3%
113
arxiv · arxiv/2604.07023 · 2026-04-09
Lightweight fine-tuning that teaches AR models to predict multiple tokens per forward pass. At 7B on Qwen2.5, MARS improves GSM8K by 4.5 points and HumanEval by 3.0 points over AR SFT while enabling 1
26.6%
114
arxiv · arxiv/2604.08468 · 2026-04-09
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
26.3%
115
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
24.3%
116
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
24.2%
117
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
23.8%
118
arxiv · arxiv/2604.07023 · 2026-04-09
Lightweight fine-tuning that teaches AR models to predict multiple tokens per forward pass. At 7B on Qwen2.5, MARS improves GSM8K by 4.5 points and HumanEval by 3.0 points over AR SFT while enabling 1
19.4%