AIME 2025 pass@1 with Python interpreter access. 100% consensus@8. Retired Feb 2026. Tool-assisted score — not directly comparable to models without tool access.
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
New 48B MoE model (2.7B active) trained on 20T tokens with novel FiberPO RL algorithm. Open-sourced on HuggingFace. Competitive with much larger models on SWE-bench Verified (62.6%), LiveCodeBench v6
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
Label-free RL via reasoning structure rewards on Qwen3-4B matches or beats ground-truth RL: AIME25 45.83% vs 46.67% oracle, MATH500 93.30%, and Minerva 61.99%.
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
Label-free RL via reasoning structure rewards on Qwen3-4B matches or beats ground-truth RL: AIME25 45.83% vs 46.67% oracle, MATH500 93.30%, and Minerva 61.99%.
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Two-phase GRPO framework: first solve, then refine. Qwen3-4B ThinkTwice achieves 60.43% AIME (vs 29.18% base) and 95.70% MATH500 after self-refinement, beating DAPO and DrGRPO.
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
Stable-OPD stabilizes on-policy distillation for small math models, boosting Qwen2.5-Math-1.5B from 28.9% to 36.1% avg accuracy and Qwen2.5-Math-7B to 47.6%, surpassing GRPO and SFT baselines on math
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
TTC-Net adds test-time compute layers to Llama-3-Instruct-7B via hardware-efficient optimal control. Acc@8 scores: MATH-500 improves from 25.0% to 52.8%, AIME25 from 0% to 5.0%. The TTC layer runs in
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
TTC-Net adds test-time compute layers to Llama-3-Instruct-7B via hardware-efficient optimal control. Acc@8 scores: MATH-500 improves from 25.0% to 52.8%, AIME25 from 0% to 5.0%. The TTC layer runs in
Evolutionary model merging produces Pareto-optimal reasoning models that reduce output length by 50%+ while matching or improving accuracy on math benchmarks.