AIME
117 models tested · Updated 2025-12-11 · Verified sources only
GPT-5.2 leads at 100.0%
1
Perfect score on math competition. No tools enabled.
100.0%
2
Anthropic · Blog/Anthropic · 2026-02-05
Perfect score on AIME 2025. Matches GPT-5.2 ceiling — AIME may be saturated for frontier models.
100.0%
3
OpenAI · Blog/OpenAI · 2025-04-16
AIME 2025 pass@1 with Python interpreter access. 100% consensus@8. Retired Feb 2026. Tool-assisted score — not directly comparable to models without tool access.
99.5%
4
Z.ai · Blog/Z.ai · 2026-06-13
Up from GLM-5.1 (95.3). Near-perfect math competition score.
99.2%
5
OpenAI · Blog/OpenAI · 2025-08-05
AIME 2025 with tool use. Edges out gpt-oss-120b (97.9%). Remarkable for a 3.6B active param model.
98.7%
6
ByteDance · Blog/ByteDance · 2026-02-14
Near-perfect AIME score. ByteDance reports gold medals on major math olympiads.
98.3%
7
Google · Leaderboard/MathArena · 2026-02-28
AIME 2024+2025 combined (60 questions). New AIME SOTA.
98.13%
8
OpenAI · Blog/OpenAI · 2025-08-05
AIME 2025 with tool use. Also scores 96.6% on AIME 2024 (tools). Apache 2.0 open-weight.
97.9%
9
StepFun · Blog/StepFun · 2026-02-12
Open-weight MoE (196B total, 11B active). Reaches 99.9 with PaCoRe parallel thinking. Best open model on AIME.
97.3%
10
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Near-frontier math; trails GPT-5.6 Sol (99.9) but matches DeepSeek V4 Pro. Open-weights 975B MoE.
97.1%
11
Alibaba · Blog/Z.ai · 2026-06-13
Strong but below GLM-5.2 (99.2) and GPT-5.5 (98.3).
97.0%
12
Google · Vals.ai · 2026-03-17
MathArena AIME 2026. 95% without tools, 100% with code execution. Top math performance.
96.7%
13
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
AIME 2026. Slightly up from K2.5 (96.1%). Close to Opus 4.6 (96.7%) but below GPT-5.4 (99.2%).
96.4%
14
AIME 2025. Strong math reasoning for a 13B-active MoE model.
96.3%
15
Moonshot AI · HuggingFace/moonshotai · 2026-01-27
1T MoE, 32B active. Open-source. Averaged over 32 runs on AIME 2025.
96.1%
16
DeepSeek · Blog/DeepSeek · 2025-12-01
AIME 2025. Deep reasoning variant. SOTA-level math, surpasses GPT-5 (94.6).
96.0%
17
Zhipu AI · MathArena · 2026-04-08
AIME 2026 competition. MathArena independent evaluation, 3rd place behind Step-3.5-Flash (96.67).
95.83%
18
Zhipu AI · NVIDIA/Zhipu Official · 2025-12-22
Open-weight 400B model. Highest AIME score among open models, competitive with GPT-5.4.
95.7%
19
Anthropic · Blog/Z.ai · 2026-06-13
Lower than GLM-5.2 (99.2) and GPT-5.5 (98.3).
95.7%
20
Alibaba · Blog/Qwen · 2026-04-02
Highest AIME score among Qwen models. Strong math reasoning, competitive with Gemini 3 Pro (96.7) and Kimi K2.5 (96.1).
95.3%
21
Zhipu AI · MathArena · 2026-04-08
AIME 2026 competition. MathArena independent evaluation, 4th place.
95.3%
22
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
AIME 2026. Strong math reasoning at 12B active params.
95.1%
23
DeepSeek · HuggingFace/deepseek-ai · 2026-04-24
Remarkable math for 13B activated params. Near Pro-level.
94.8%
24
OpenAI · Blog/OpenAI · 2025-08-07
AIME 2025 without tools. Strong but not perfect — Claude Opus 4.6 and GPT-5.2 both hit 100%.
94.6%
25
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
94.2%
26
DeepSeek · Blog/DeepSeek · 2025-12-01
AIME 2026. 685B MoE params with DeepSeek Sparse Attention. Top-tier math reasoning.
94.17%
27
Xiaomi · arxiv/2601.02780 · 2026-01-06
15B active MoE. AIME 2025 score, near frontier level.
94.1%
28
Alibaba · HuggingFace/Qwen · 2026-04-22
Beats Gemma 4 31B by 5pp on AIME 2026.
94.1%
29
Anthropic · Blog/Anthropic · 2026-02-17
Advanced math reasoning. Anthropic official benchmarks page.
94.0%
30
xAI · Artificial Analysis · 2025-07-09
Joint highest AIME 2024 at time of release. Third-party verified.
94.0%
31
DeepSeek · Paper/DeepSeek-V3.2 · 2025-12-01
AIME 2025, Pass@1. Open-weight, 685B MoE. Matches GPT-5 on reasoning.
93.1%
32
ByteDance · Blog/ByteDance · 2026-03-10
Mid-tier model in Seed 2.0 family. Handles 95% of enterprise workloads at half the Pro cost.
93.0%
33
Zhipu AI · Paper/Zhipu AI (arxiv:2602.15763) · 2026-02-11
AIME 2026 I. On par with Kimi K2.5 (92.5) and DeepSeek V3.2 (92.7).
92.7%
34
OpenAI · Blog/OpenAI · 2025-04-16
Without tools (closed-book). With tools achieves 99.5%. Best-performing small reasoning model on AIME.
92.7%
35
Alibaba · HuggingFace/Qwen · 2026-04-16
AIME 2026 score. Matches GLM-5 and beats Gemma 4 31B (89.2) as an open MoE.
92.7%
36
LG AI Research · HuggingFace/LGAI-EXAONE · 2026-04-14
AIME 2026 variant. Competitive with Qwen3.5 27B (90.8) despite being a VLM. Dense 33B params.
92.6%
37
Alibaba · HuggingFace/Qwen · 2026-03-02
MathArena AIME 2026. A 9B model matching frontier math scores. Remarkable reasoning-per-parameter.
92.5%
38
NVIDIA · arxiv/2603.19220 · 2026-03-19
Second open-weight LLM to achieve IMO/IOI/ICPC Gold; 30B MoE with only 3B active params matching frontier models on math reasoning.
92.4%
39
Zhipu AI · HuggingFace/Zhipu · 2026-01-15
30B-A3B MoE with only 3.6B active params. Strong math reasoning for a lightweight model.
91.6%
40
Alibaba · HuggingFace/Qwen · 2026-02-16
AIME 2026. Strong math reasoning from open-weight MoE model.
91.3%
41
PrismML · Whitepaper/PrismML · 2026-07-14
Ternary variant, AIME25. Conventional IQ2_XXS collapses to 66.67 here; ternary Bonsai retains 90.84 at half the footprint.
90.84%
42
Alibaba · HuggingFace/Alibaba · 2026-02-20
AIME 2026. Exceptional math reasoning for a 27B model — near frontier-class.
90.83%
43
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
90.1%
44
arxiv · arxiv/2604.06753 · 2026-04-08
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
90.0%
45
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
90.0%
46
Google · X/@o_mega___ · 2026-04-02
Independent user corroborates Google's AIME score; 0 engagement
89.2%
47
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
89.2%
48
NVIDIA · HuggingFace/NVIDIA · 2025-12-15
Without tools. Hybrid Mamba-Transformer MoE, 3.3x throughput vs Qwen3-30B on H200.
89.1%
49
PrismML · Whitepaper/PrismML · 2026-07-14
1-bit variant, AIME25. Binary weights retain 88.75 vs 93.29 FP16 — reasoning is the most compression-robust capability, per PrismML.
88.75%
50
Google · Model Card/Google · 2026-04-02
MoE with only 3.8B active params outperforms many larger models.
88.3%
51
Sarvam AI · HuggingFace/sarvamai · 2026-03-06
India's first domestically-trained 105B model. MoE with 10.3B active params. Apache 2.0.
88.3%
52
PrismML · Whitepaper/PrismML · 2026-07-14
Ternary variant, AIME26. The benchmark where conventional sub-4-bit breaks hardest (IQ2_XXS: 57.50). Bonsai holds 87.5 at 1.71 bits/weight.
87.5%
53
PrismML · Whitepaper/PrismML · 2026-07-14
1-bit variant, AIME26. Both Bonsai variants hold AIME above 87 at sub-2-bit; the exact capability conventional sub-4-bit loses.
87.08%
54
MiniMax · Blog/MiniMax · 2026-02-12
Competitive math reasoning. Behind Claude Opus 4.5 (91.0) but ahead of Sonnet 4.5 (88.0).
86.3%
55
arxiv · arxiv/2604.08299 · 2026-04-09
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
83.33%
56
arxiv · arxiv/2604.08299 · 2026-04-09
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
83.33%
57
Alibaba Cloud · arxiv/2601.09088 · 2026-01-15
4B model beating 32B and 253B models on AIME 2025. Trained on 448K samples with Distribution-Aligned Sequence Distillation.
83.3%
58
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
82.0%
59
Microsoft · HuggingFace/Microsoft · 2025-04-30
SFT+RL variant of Phi-4-reasoning. 14B model hitting 81%+ on AIME 2024 — remarkable for size.
81.3%
60
Anthropic · Blog/Anthropic · 2025-10-15
Strong math reasoning for a Haiku-tier model. Extended thinking with 128K budget.
80.7%
61
ServiceNow · arxiv/2604.02007 · 2026-04-02
15B open-weight model. RL post-training achieves strong accuracy with 30-50% shorter reasoning traces. Trained on 5 domains with adaptive sampling.
78.3%
62
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
77.8%
63
Google DeepMind · HuggingFace/google-gemma-4-12b-it · 2026-05-23
AIME 2026 no tools. Beats Gemma 3 27B (20.8) decisively.
77.5%
64
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
77.3%
65
arxiv · arxiv/2603.25284 · 2026-03-26
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
76.67%
66
arxiv · arxiv/2603.25284 · 2026-03-26
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
76.67%
67
Microsoft · HuggingFace/Microsoft · 2025-04-30
14B dense model. Matches DeepSeek-R1 (671B) on AIME despite 47x fewer params.
75.3%
68
arxiv · arxiv/2603.25284 · 2026-03-26
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
73.33%
69
arxiv · arxiv/2604.03044 · 2026-04-03
New 48B MoE model (2.7B active) trained on 20T tokens with novel FiberPO RL algorithm. Open-sourced on HuggingFace. Competitive with much larger models on SWE-bench Verified (62.6%), LiveCodeBench v6
72.9%
70
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
70.0%
71
arxiv · arxiv/2603.25284 · 2026-03-26
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
70.0%
72
arxiv · arxiv/2604.05424 · 2026-04-06
MCTS with metacognitive reflection improves reasoning. On Qwen3-30B-A3B, PRISM-MCTS achieves 65.15% on GPQA Diamond and 93.28% on MATH500.
63.33%
73
Qwen · arxiv/2604.05355 · 2026-04-07
AIME24 baseline; ETR boosts to 73.3. Strong AIME24 performance for 8B dense model.
63.3%
74
arxiv · arxiv/2604.01591 · 2026-04-02
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
60.43%
75
arxiv · arxiv/2604.06753 · 2026-04-08
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
58.3%
76
arxiv · arxiv/2604.05267 · 2026-04-07
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
56.7%
77
arxiv · arxiv/2604.05424 · 2026-04-06
MCTS with metacognitive reflection improves reasoning. On Qwen3-30B-A3B, PRISM-MCTS achieves 65.15% on GPQA Diamond and 93.28% on MATH500.
53.33%
78
arxiv · arxiv/2604.05424 · 2026-04-06
MCTS with metacognitive reflection improves reasoning. On Qwen3-30B-A3B, PRISM-MCTS achieves 65.15% on GPQA Diamond and 93.28% on MATH500.
53.33%
79
arxiv · arxiv/2604.08299 · 2026-04-09
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
53.33%
80
Qwen · arxiv/2604.05355 · 2026-04-07
AIME24 baseline; ETR boosts to 73.3. Compact 4B model with strong math reasoning via efficient CoT optimization.
53.3%
81
arxiv · arxiv/2604.08468 · 2026-04-09
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
50.0%
82
arxiv · arxiv/2604.05424 · 2026-04-06
MCTS with metacognitive reflection improves reasoning. On Qwen3-30B-A3B, PRISM-MCTS achieves 65.15% on GPQA Diamond and 93.28% on MATH500.
48.67%
83
arxiv · arxiv/2604.06465 · 2026-04-07
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
46.7%
84
arxiv · arxiv/2604.08299 · 2026-04-09
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
46.67%
85
arxiv · arxiv/2603.27977 · 2026-03-30
Label-free RL via reasoning structure rewards on Qwen3-4B matches or beats ground-truth RL: AIME25 45.83% vs 46.67% oracle, MATH500 93.30%, and Minerva 61.99%.
45.83%
86
arxiv · arxiv/2604.01591 · 2026-04-02
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
44.11%
87
arxiv · arxiv/2604.06465 · 2026-04-07
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
43.3%
88
Google · Google Model Card · 2026-04-02
Impressive for a 4B model running on T4 GPUs. MoE architecture.
42.5%
89
arxiv · arxiv/2604.06753 · 2026-04-08
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
40.0%
90
arxiv · arxiv/2604.01591 · 2026-04-02
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
39.24%
91
Google · Model Card/Google · 2026-04-02
Matches Gemma 3 27B despite being a fraction of the size. Built-in reasoning capability.
37.5%
92
arxiv · arxiv/2604.06465 · 2026-04-07
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
36.7%
93
arxiv · arxiv/2604.08468 · 2026-04-09
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
36.7%
94
arxiv · arxiv/2604.06465 · 2026-04-07
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
33.3%
95
arxiv · arxiv/2604.06465 · 2026-04-07
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
33.3%
96
arxiv · arxiv/2604.01591 · 2026-04-02
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
32.81%
97
arxiv · arxiv/2603.27977 · 2026-03-30
Label-free RL via reasoning structure rewards on Qwen3-4B matches or beats ground-truth RL: AIME25 45.83% vs 46.67% oracle, MATH500 93.30%, and Minerva 61.99%.
31.67%
98
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
29.6%
99
arxiv · arxiv/2604.01591 · 2026-04-02
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
29.18%
100
arxiv · arxiv/2604.06465 · 2026-04-07
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
26.7%
101
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
26.6%
102
arxiv · arxiv/2604.08527 · 2026-04-09
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
25.1%
103
arxiv · arxiv/2604.08527 · 2026-04-09
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
24.7%
104
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
23.7%
105
arxiv · arxiv/2604.08468 · 2026-04-09
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
23.3%
106
arxiv · arxiv/2604.08527 · 2026-04-09
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
22.2%
107
arxiv · arxiv/2604.08468 · 2026-04-09
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
16.7%
108
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
16.3%
109
arxiv · arxiv/2604.08527 · 2026-04-09
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
13.8%
110
arxiv · arxiv/2604.08527 · 2026-04-09
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
11.8%
111
arxiv · arxiv/2604.08527 · 2026-04-09
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
11.7%
112
arxiv · arxiv/2603.25284 · 2026-03-26
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
10.0%
113
arxiv · arxiv/2604.06753 · 2026-04-08
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
10.0%
114
arxiv · arxiv/2603.09221 · 2026-03-10
TTC-Net adds test-time compute layers to Llama-3-Instruct-7B via hardware-efficient optimal control. Acc@8 scores: MATH-500 improves from 25.0% to 52.8%, AIME25 from 0% to 5.0%. The TTC layer runs in
5.0%
115
arxiv · arxiv/2603.05357 · 2026-03-05
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
2.9%
116
arxiv · arxiv/2603.09221 · 2026-03-10
TTC-Net adds test-time compute layers to Llama-3-Instruct-7B via hardware-efficient optimal control. Acc@8 scores: MATH-500 improves from 25.0% to 52.8%, AIME25 from 0% to 5.0%. The TTC layer runs in
0.0%
117
arxiv · arxiv/2604.06465 · 2026-04-07
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
0.0%