SWE-bench Verified
86 models tested · Updated 2026-04-07 · Verified sources only
Claude Mythos Preview leads at 93.9%
1
Anthropic · Blog/Anthropic · 2026-04-07
New SOTA. 13.1pp above Opus 4.6 (80.8%). Largest single-generation coding jump in Anthropic history.
93.9%
2
Anthropic · Blog/Anthropic · 2026-04-16
Up 6.8pp from Opus 4.6. New SOTA on SWE-bench Verified, surpassing Claude Mythos Preview.
87.6%
3
Anthropic · Blog/Anthropic · 2026-02-05
Updated official score with prompt modification, averaged over 25 trials. Previous entry was 80.8%.
81.42%
4
Anthropic · Blog/Anthropic · 2025-11-24
First model to exceed 80% on SWE-bench Verified. 25 trials avg.
80.9%
5
Google · Official/Google DeepMind · 2026-02-19
Agentic coding benchmark. Verified from official Gemini page.
80.6%
6
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning. Matches Gemini 3.1 Pro, below Opus 4.7 (87.6).
80.6%
7
MiniMax · Blog/MiniMax · 2026-02-12
MoE 230B params, 10B active. Matches Claude Opus 4.6 speed on SWE-Bench eval.
80.2%
8
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
Open-weight model matching Opus 4.6 (80.8%) and Gemini 3.1 Pro (80.6%) on SWE-bench Verified.
80.2%
9
OpenAI · OpenAI — Introducing GPT-5.2 · 2025-12-11
Multi-language bug fixing. SWE-bench Pro score: 55.6% (new SOTA).
80.0%
10
OpenAI · Blog/OpenAI · 2026-02-05
OpenAI coding-focused model. Also scores 56.8% on SWE-bench Pro (harder multilingual variant).
80.0%
11
Z.ai · Blog/ThinkingMachines · 2026-07-15
From TML Inkling blog comparison table. GLM-5.2 official is 80.0.
80.0%
12
Anthropic · Blog/Anthropic · 2026-02-17
79.6% on SWE-bench Verified, 1.2 points behind Opus 4.6, surpasses Sonnet 4.5 by 2.4 points
79.6%
13
Kuaishou · arxiv/2603.27703 · 2026-03-29
Near Claude Opus 4.6 (80.8%) using Claude Code scaffold. Specialize-then-Unify paradigm with MCLA for MoE RL training.
79.6%
14
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning, 13B activated. Impressive for small model, near V4 Pro (80.6).
79.0%
15
Alibaba · Blog/Qwen · 2026-04-02
Qwen 3.6 Plus scores 78.8% on SWE-bench Verified
78.8%
16
Google · Blog/Google · 2025-12-17
Distilled model outperforms Gemini 3 Pro (76.2%) on agentic coding. Less than a quarter the cost of Pro.
78.0%
17
MiniMax · Blog/MiniMax · 2026-03-18
Self-improving MoE model activating only 10B params. Scores below M2.5 (80.2%) on this benchmark — may use different eval config or scaffold.
78.0%
18
Tencent · HuggingFace/tencent-Hy3-modelcard · 2026-07-06
Open-weight 295B MoE (21B active). Official Tencent score, not the 74.4 rumour. Beats DeepSeek V4 Pro (80.6 is higher actually - check). Strong agentic model.
78.0%
19
Zhipu · Zhipu — GLM-5 Developer Docs · 2026-02-11
Leading open-model score. Only 3 points behind Opus 4.6. Terminal-Bench 2.0: 56.2%.
77.8%
20
Zhipu · X/@grok · 2026-03-27
Grok bot relayed Zhipu GLM-5.1 score; no direct link to paper or official source
77.8%
21
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Self-reported via bash-only harness. Competitive with Kimi K2.5 (76.8) but below GLM-5.2 (80.0).
77.6%
22
Meta · Blog/Meta · 2026-04-08
First model from Meta Superintelligence Labs. Competitive with Gemini 3.1 Pro on agentic coding.
77.4%
23
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
Matches Inkling on SWE-bench Verified at a third the active params.
77.4%
24
Anthropic · Blog/Anthropic · 2025-09-29
77.2% standard, 78.2% with 1M context, 82.0% with parallel compute. SOTA at time of release.
77.2%
25
OpenAI · Vals.ai · 2026-03-05
Independent Vals.ai evaluation. OpenAI dropped SWE-bench Verified in favor of SWE-bench Pro (57.7%).
77.2%
26
Alibaba · HuggingFace/Qwen · 2026-04-22
Best open 27B dense model on SWE-bench. Outperforms Gemma 4 31B by 25pp.
77.2%
27
Moonshot AI · HuggingFace/moonshotai · 2026-01-27
1T MoE open-source. Competitive with Claude Sonnet 4.6 (79.6). Uses internal eval framework with bash/createfile tools.
76.8%
28
ByteDance · Blog/ByteDance · 2026-02-14
Competitive with Claude Sonnet 4.5 and DeepSeek V3.2 on coding tasks.
76.5%
29
StepFun · Blog/StepFun · 2026-05-29
SWE-bench Verified. Beats Step 3.5 Flash (74.4).
76.5%
30
Alibaba · HuggingFace/Qwen · 2026-02-16
Strong open-weight coding performance. 17B active params, MoE architecture.
76.4%
31
OpenAI · Blog/OpenAI · 2025-08-07
Below GPT-5.4 (77.2%) and Claude Opus 4.6 (81.4%). Coding was not GPT-5 main strength.
74.9%
32
Fujitsu Research · Blog/Fujitsu Research · 2026-04-08
Third-party harness engineering by Fujitsu. Test-time scaling with 8 candidate patches, no fine-tuning. Highest score among <229B param models.
74.8%
33
StepFun · Blog/StepFun · 2026-02-12
Open-weight MoE with only 11B active params. Competitive with proprietary mid-tier models on coding.
74.4%
34
xAI · X/@grok · 2026-03-31
Top tier with Claude/GPT. Fast inference + huge context.
74.0%
35
Zhipu AI · NVIDIA/Zhipu Official · 2025-12-22
Open-weight, 400B params. Strong coding for an open model but behind GLM-5/5.1 successors.
73.8%
36
ByteDance · Blog/ByteDance · 2026-03-10
Near Pro level (76.5) at roughly half the cost.
73.5%
37
Xiaomi · arxiv/2601.02780 · 2026-01-06
15B active MoE model. Competitive with models 10-20x larger. Uses Multi-Teacher On-Policy Distillation.
73.4%
38
Alibaba · HuggingFace/Qwen · 2026-04-16
Open-weight MoE with only 3B active params. Beats Gemma 4 31B (52.0) by +21.4pts on same benchmark.
73.4%
39
Anthropic · Blog/Anthropic · 2025-10-15
Anthropic first Haiku with extended thinking. Matches Sonnet 4 on coding at 1/3 cost.
73.3%
40
DeepSeek · Paper/DeepSeek (arxiv:2512.02556) · 2025-12-01
Primary score from DeepSeek technical report. Robustness tests across frameworks range 72-74%. Trails Claude Opus 4.6 (80.8%) and GPT-5.3 Codex (80.0%).
73.1%
41
xAI · X/@grok · 2026-03-31
Grok bot stated ~70-75% range; midpoint used; no official xAI source linked
72.5%
42
Alibaba · Model Card/Qwen · 2026-02-24
Open-weight 27B dense model matching GPT-5 mini. Released Feb 2026.
72.4%
43
Mistral · Mistral — Devstral 2 and Vibe CLI · 2025-12-09
123B dense transformer, 256k context. Best open-weight SWE-bench score. 7x more cost-efficient than Claude Sonnet.
72.2%
44
OpenAI · Blog/OpenAI · 2025-04-16
Solid coding ability. Behind frontier models (Claude Opus 4.6 at 81.4%, Gemini 3.1 Pro at 80.6%).
71.7%
45
Qwen · arxiv/2603.00729 · 2026-02-28
OpenHands agent score; SWE-Agent gets 70.6. 3B active params competitive with 35B+ models.
71.3%
46
arxiv · arxiv/2604.01496 · 2026-04-02
SWE-Hero-32B achieves 62.2% on SWE-bench Verified and 44.1% on SWE-bench Lite with execution-based fine-tuning. Open-weight 32B model competitive with MiniMax-M2.5 (80.2% Verified).
71.3%
47
NVIDIA · Blog/ThinkingMachines · 2026-07-15
From TML Inkling blog. NVIDIA's Nemotron 3 Ultra.
70.7%
48
Alibaba · Blog/Alibaba · 2026-02-03
SWE-Agent score. 80B MoE, only 3B active params. Coding specialist matching models 10-20x larger.
70.6%
49
Kuaishou · arxiv/2604.03144 · 2026-04-03
32B open-weight model with ECoT reasoning traces and industrial code world model. Top open-weight LiveCodeBench result.
70.4%
50
Meta · Blog/Meta · 2026-04-05
17B MoE with 128 experts. Strong coding performance for open-weight model, competitive with GPT-4o class.
70.3%
51
Qwen · arxiv/2604.01496 · 2026-04-02
MoE (480B total, 35B active) code model with OpenHands agent scaffold. Without execution: 59.4%.
69.6%
52
arxiv · arxiv/2604.01496 · 2026-04-02
SWE-Hero-32B achieves 62.2% on SWE-bench Verified and 44.1% on SWE-bench Lite with execution-based fine-tuning. Open-weight 32B model competitive with MiniMax-M2.5 (80.2% Verified).
69.6%
53
OpenAI · Blog/OpenAI · 2025-04-16
Cost-efficient reasoning model. Behind o3 (69.1%) but far ahead of o1 (48.9%) and o3-mini (49.3%).
68.1%
54
Mistral · Blog/Mistral · 2025-12-09
24B open-weight coding model. Runs locally on RTX 4090. Apache 2.0 license.
68.0%
55
arxiv · arxiv/2604.05407 · 2026-04-07
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
67.2%
56
arxiv · arxiv/2604.05407 · 2026-04-07
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
66.2%
57
arxiv · arxiv/2604.05407 · 2026-04-07
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
66.0%
58
Evaluated with mini-swe-agent-v2. Behind closed-source leaders but competitive for open-weight.
63.2%
59
arxiv · arxiv/2604.03044 · 2026-04-03
New 48B MoE model (2.7B active) trained on 20T tokens with novel FiberPO RL algorithm. Open-sourced on HuggingFace. Competitive with much larger models on SWE-bench Verified (62.6%), LiveCodeBench v6
62.6%
60
OpenAI · HuggingFace/openai/gpt-oss-120b · 2025-08-05
OpenAI's first open-weight model at 116.8B params (5.1B active). Apache 2.0 license. High reasoning effort.
62.4%
61
Unknown · arxiv/2604.01496 · 2026-04-02
32B open-source SWE agent. Top open-source model on SWE-bench Verified (excludes Git hacking results).
62.4%
62
NVIDIA · arxiv/2604.01496 · 2026-04-02
Open-weight 32B model distilled from Qwen3-Coder-480B. Two-stage SFT: execution-free then execution-backed. New SOTA for open-source 32B-class models on SWE-bench.
62.2%
63
arxiv · arxiv/2604.01496 · 2026-04-02
SWE-Hero-32B achieves 62.2% on SWE-bench Verified and 44.1% on SWE-bench Lite with execution-based fine-tuning. Open-weight 32B model competitive with MiniMax-M2.5 (80.2% Verified).
62.2%
64
arxiv · arxiv/2604.05407 · 2026-04-07
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
62.0%
65
Unknown · arxiv/2604.01496 · 2026-04-02
32B model trained with SFT+RL via R2E-Gym. Second-best open-source 32B-class model on SWE-bench.
61.4%
66
arxiv · arxiv/2604.05407 · 2026-04-07
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
61.2%
67
NVIDIA · arxiv/2604.01496 · 2026-04-02
Mid-size SWE-Hero variant competitive with much larger models. Distilled from Qwen3-Coder-480B.
60.8%
68
OpenAI · HuggingFace/openai/gpt-oss-20b · 2025-08-05
20.9B total params (3.6B active). Runs on 16GB consumer hardware. Apache 2.0. Impressive for its size — near gpt-oss-120b SWE-bench score.
60.7%
69
NVIDIA · Blog/NVIDIA · 2026-03-11
Highest open-weight SWE-bench Verified score. Evaluated with OpenHands scaffold.
60.47%
70
OpenAI · arxiv/2604.05407 · 2026-04-07
Mid-size GPT-5 variant. 60.4% with SWE-Agent, improves to 62.0% with structured code actions.
60.4%
71
Moonshot AI · HuggingFace/Moonshot AI · 2025-09-27
Open-weight 72B coding LLM, RL-trained in Docker containers (reward only when full test suite passes). SOTA among open-source workflow approaches at release.
60.4%
72
Zhipu AI · HuggingFace/Zhipu · 2026-01-15
Lightweight MoE flash variant. Solid for its size class.
59.2%
73
xAI · Leaderboard/vals.ai · 2026-03-20
Independent vals.ai test. Significant gap from xAI self-reported 72-75%, highlighting scaffold sensitivity.
58.6%
74
NVIDIA · arxiv/2604.01496 · 2026-04-02
Smallest SWE-Hero variant. Outperforms same-size open-source agents by significant margin.
52.7%
75
Google · HuggingFace/Qwen · 2026-04-02
Dense 31B model. Lagging behind MoE competitors like Qwen 3.6 35B-A3B (73.4) on agentic coding.
52.0%
76
NVIDIA · arxiv/2603.19220 · 2026-03-19
SWE-bench Verified via OpenHands agent; lower than Qwen3.5-35B (69.2) showing SWE tasks still challenge compact MoE models.
50.2%
77
Sarvam AI · HuggingFace/sarvamai · 2026-03-06
India's first domestically-trained 105B model. MoE with 10.3B active params. Apache 2.0.
45.0%
78
Huawei · arxiv/2604.00824 · 2026-04-01
30B MoE (3B active) fine-tuned with STITCH trajectory filtering. 63% relative improvement over base (26.6%).
43.4%
79
SWE-Protege Team · arxiv/2602.22124 · 2026-02-25
+25.4% over prior SLM SOTA on SWE-bench Verified. Expert-protégé collaboration framework enables 7B model to selectively seek guidance while remaining sole decision-maker.
42.4%
80
Zhipu AI · arxiv/2604.00824 · 2026-04-01
MoE model (106B total, 12B active). Baseline with SWE-Agent; improves to 48.4% with STITCH trajectory filtering.
42.2%
81
arxiv · arxiv/2604.05407 · 2026-04-07
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
40.4%
82
OpenAI · arxiv/2604.05407 · 2026-04-07
Smallest GPT-5 variant. Struggles with SWE-agent workflows (46.6% empty-patch rate), but improves to 40.4% with structured code actions.
19.6%
83
Google · HuggingFace/Qwen · 2026-04-18
MoE model with 4B active params. Significantly behind Qwen 3.6 35B-A3B on agentic coding.
17.4%
84
arxiv · arxiv/2604.05407 · 2026-04-07
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
16.0%
85
arxiv · arxiv/2604.05407 · 2026-04-07
Reframes codebase as structured AST action space. GPT-5-nano + CodeStruct jumps from 19.6% to 40.4% on SWE-bench Verified (+20.8). GPT-5 + CodeStruct hits 67.2%. Reduces errors by up to 87.8%.
14.8%
86
Qwen · arxiv/2604.05407 · 2026-04-07
Small open-weight model baseline on SWE-Agent. Too small for effective autonomous SWE without specialized fine-tuning.
13.2%