Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
New 48B MoE model (2.7B active) trained on 20T tokens with novel FiberPO RL algorithm. Open-sourced on HuggingFace. Competitive with much larger models on SWE-bench Verified (62.6%), LiveCodeBench v6
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
Domain Steering MoE (DSMoE): training-free framework that activates domain-specific experts in sparse MoE LLMs at zero additional inference cost. Key finding: GPT-OSS-20B (3.6B active params) with DSM
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Selective latent reasoning that activates continuous reasoning only at uncertain tokens via entropy threshold, achieving higher accuracy than standard CoT with fewer tokens. Qwen3-8B + SeLaR matches Q
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Lightweight fine-tuning that teaches AR models to predict multiple tokens per forward pass. At 7B on Qwen2.5, MARS improves GSM8K by 4.5 points and HumanEval by 3.0 points over AR SFT while enabling 1
Test-time variational synthesis improves RLVR for small models without labeled data. Qwen3-4B + TTVS achieves 90.3% MATH500 and 48.9% GPQA, outperforming baseline by +26.3 and +21.7 points respectivel
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Self-curriculum test-time RL method that improves small models substantially. Qwen3-4B-Base with DiSCTT reaches 81.3% MMLU and 75.2% MATH-500, surpassing Qwen2.5-7B-Instruct base (76.2% MMLU, 58.8% MA
Lightweight fine-tuning that teaches AR models to predict multiple tokens per forward pass. At 7B on Qwen2.5, MARS improves GSM8K by 4.5 points and HumanEval by 3.0 points over AR SFT while enabling 1