Open-weight 295B MoE (21B active). Official Tencent score, not the 74.4 rumour. Beats DeepSeek V4 Pro (80.6 is higher actually - check). Strong agentic model.
Primary score from DeepSeek technical report. Robustness tests across frameworks range 72-74%. Trails Claude Opus 4.6 (80.8%) and GPT-5.3 Codex (80.0%).
SWE-Hero-32B achieves 62.2% on SWE-bench Verified and 44.1% on SWE-bench Lite with execution-based fine-tuning. Open-weight 32B model competitive with MiniMax-M2.5 (80.2% Verified).
SWE-Hero-32B achieves 62.2% on SWE-bench Verified and 44.1% on SWE-bench Lite with execution-based fine-tuning. Open-weight 32B model competitive with MiniMax-M2.5 (80.2% Verified).
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
New 48B MoE model (2.7B active) trained on 20T tokens with novel FiberPO RL algorithm. Open-sourced on HuggingFace. Competitive with much larger models on SWE-bench Verified (62.6%), LiveCodeBench v6
Open-weight 32B model distilled from Qwen3-Coder-480B. Two-stage SFT: execution-free then execution-backed. New SOTA for open-source 32B-class models on SWE-bench.
SWE-Hero-32B achieves 62.2% on SWE-bench Verified and 44.1% on SWE-bench Lite with execution-based fine-tuning. Open-weight 32B model competitive with MiniMax-M2.5 (80.2% Verified).
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
Open-weight 72B coding LLM, RL-trained in Docker containers (reward only when full test suite passes). SOTA among open-source workflow approaches at release.
+25.4% over prior SLM SOTA on SWE-bench Verified. Expert-protégé collaboration framework enables 7B model to selectively seek guidance while remaining sole decision-maker.
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.