New 48B MoE model (2.7B active) trained on 20T tokens with novel FiberPO RL algorithm. Open-sourced on HuggingFace. Competitive with much larger models on SWE-bench Verified (62.6%), LiveCodeBench v6
Aggressive parallel decoding for diffusion LLMs that preserves accuracy while achieving 2x+ throughput gains over LLaDA-2.0-mini. At low parallelism, DMax-Math improves MATH500 from 75.8% to 78.0% and
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
Lightweight fine-tuning that teaches AR models to predict multiple tokens per forward pass. At 7B on Qwen2.5, MARS improves GSM8K by 4.5 points and HumanEval by 3.0 points over AR SFT while enabling 1
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
SliderQuant preserves accuracy in 4-bit quantized LLMs. DeepSeek-R1-Distill-Qwen-14B at W4A16 retains 94.6% MATH and 91.35% GSM8K vs full-precision. Even W2A16 keeps 29.4% MATH vs 0% for OmniQuant. Qw
Lightweight fine-tuning that teaches AR models to predict multiple tokens per forward pass. At 7B on Qwen2.5, MARS improves GSM8K by 4.5 points and HumanEval by 3.0 points over AR SFT while enabling 1
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.