SWE-bench Pro
43 models tested · Updated 2026-06-09 · Verified sources only
Claude Fable 5 leads at 80.3%
1
Anthropic · Blog/Anthropic · 2026-06-09
New SOTA. 21.7 points above GPT-5.5 (58.6%). Agentic coding at scale.
80.3%
2
Anthropic · Blog/OpenAI · 2026-06-09
Unsafeguarded Mythos-class model via Project Glasswing. Same underlying model as Fable 5 with cyber safeguards lifted.
80.3%
3
Anthropic · Blog/Anthropic · 2026-04-07
24.4pp above Opus 4.6 (53.4%). Massive jump on harder coding eval.
77.8%
4
Anthropic · Blog/Z.ai · 2026-06-13
Second only to Fable 5 (80.3). Beats GLM-5.2 (62.1).
69.2%
5
xAI · Blog/xAI · 2026-07-08
Third on SWE-bench Pro behind Fable 5 (80.4) and Opus 4.8 (69.2). Notable for token efficiency: 4.2x fewer output tokens than Opus 4.8.
64.7%
6
OpenAI · Blog/OpenAI · 2026-07-09
Beats GPT-5.5 (59.4%); trails Claude Mythos 5 (80.3%).
64.6%
7
Anthropic · Blog/Anthropic · 2026-04-16
Up 10.9pp from Opus 4.6. Largest single-version jump on SWE-bench Pro.
64.3%
8
OpenAI · Blog/OpenAI · 2026-07-09
Balanced-tier Terra beats GPT-5.5 (59.4%) at lower cost.
63.4%
9
OpenAI · Blog/OpenAI · 2026-07-09
Cheapest tier Luna beats GPT-5.5 (59.4%) on SWE-bench Pro.
62.7%
10
Z.ai · Blog/Z.ai · 2026-06-13
Strongest open-source score. Closes gap to Opus 4.8 (69.2). Up from GLM-5.1 (58.4).
62.1%
11
Alibaba · Blog/Z.ai · 2026-06-13
Above GPT-5.5 (58.6), below GLM-5.2 (62.1).
60.6%
12
MiniMax · Blog/Z.ai · 2026-06-13
Above GPT-5.5 (58.6). Competitive with GLM-5.2 (62.1).
59.0%
13
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
Open-weight SOTA on SWE-bench Pro. 1T/32B MoE matching GPT-5.4 at 57.7%.
58.6%
14
OpenAI · OpenAI Blog · 2026-04-23
Matches Kimi K2.6 SOTA. Uses fewer tokens than GPT-5.4.
58.6%
15
OpenAI · OpenAI Blog · 2026-04-23
Same as GPT-5.5, listed separately in Pro column.
58.6%
16
Z.ai · X/@Zai_org · 2026-04-07
#1 open source, #3 globally. Beats GPT-5.4 (57.7) and Claude Opus 4.6 (57.3). First Chinese model to top SWE-bench Pro. Zero Nvidia hardware.
58.4%
17
Tencent · HuggingFace/tencent-Hy3-modelcard · 2026-07-06
SWE-bench Pro. Competitive with DeepSeek V4 Pro (55.4).
57.9%
18
OpenAI · Blog/OpenAI · 2026-03-05
Harder coding benchmark variant. GPT-5.4 splits coding, reasoning, and computer use into one unified model.
57.7%
19
Anthropic · Z.ai/GLM-5.1-blog · 2026-04-08
Cited alongside GPT-5.4 (57.7) and GLM-5.1 (58.4) in multiple independent comparisons. Opus 4.6 was previous SOTA before GLM-5.1.
57.3%
20
Alibaba · X/@Alibaba_Qwen · 2026-04-20
Barely edges past Opus 4.5 (57.1). Top coding model from China. Proprietary, not open-weight.
57.3%
21
OpenAI · Web/OpenAI · 2026-03-05
Leads SWE-bench Pro. 25% faster than GPT-5.2 Codex. Covers 1,865 multi-language tasks.
56.8%
22
StepFun · Blog/StepFun · 2026-05-29
SWE-bench Pro. 2nd place among open models, behind DeepSeek V4 Pro (55.4) is actually ahead.
56.3%
23
MiniMax · Blog/MiniMax · 2026-03-18
Matches GPT-5.3 Codex level. Nearly approaches Opus best on SWE-Pro.
56.2%
24
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning. Below GPT-5.5 (58.6) and Kimi K2.6 (58.6).
55.4%
25
Google DeepMind · Blog/Google DeepMind · 2026-06-10
Public set. Below Opus 4.7 (64.3). Up from Gemini 3 Flash (49.6).
55.1%
26
OpenAI · Blog/OpenAI · 2026-03-17
Smallest GPT-5.4 variant. 2x faster than GPT-5 mini at similar coding quality.
54.4%
27
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Open-weights MoE; trails GLM-5.2 (62.1) and Fable 5 (80.0) on harder SWE variant.
54.3%
28
Google · Blog/Google DeepMind · 2026-02-19
Competitive with GPT-5.4 (57.7%). Behind Claude Mythos Preview (77.8%).
54.2%
29
Alibaba · HuggingFace/Qwen · 2026-04-22
Competitive with Qwen 3.5 397B MoE on agent coding.
53.5%
30
Anthropic · arxiv/Mythos-System-Card · 2026-04-07
Scored by Anthropic using Scale AI variant. Lower than GPT-5.4 (57.7%) and Gemini 3.1 Pro (54.2%) on this harder variant.
53.4%
31
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
SWE-bench Pro Public.
53.2%
32
DeepSeek · HuggingFace/deepseek-ai · 2026-04-24
Impressive for 13B activated — competitive with much larger models.
52.6%
33
OpenAI · Blog/OpenAI · 2026-03-17
Fastest, cheapest GPT-5.4 variant. Designed for classification, data extraction, and coding subagents.
52.4%
34
Meta · Blog/Meta · 2026-04-08
Trails Claude Mythos (77.8%) significantly but competitive with other frontier models.
52.0%
35
Moonshot AI · HuggingFace/Moonshot · 2026-01-29
Internal evaluation framework with minimal tool set. 1T param MoE, 32B active.
50.7%
36
Alibaba · HuggingFace/Qwen · 2026-04-16
3B active params MoE beating Gemma 4 31B (35.7%) on SWE-bench Pro.
49.5%
37
ByteDance · Blog/ByteDance · 2026-02-14
Flagship coding model from ByteDance Seed team.
46.9%
38
NVIDIA · Blog/ThinkingMachines · 2026-07-15
From TML Inkling blog. SWE-bench Pro Public.
46.4%
39
ByteDance · Blog/ByteDance · 2026-03-10
Nearly matches Pro variant (46.9) on harder SWE subset.
46.0%
40
Alibaba · Blog/Alibaba · 2026-02-03
SWE-Agent score. Competitive with much larger models on harder SWE subset.
44.3%
41
Qwen · arxiv/2603.00729 · 2026-02-28
3B active params matching models with 10x more active compute on SWE-bench Pro.
42.7%
42
Google · HuggingFace/Qwen · 2026-04-02
Struggles on harder SWE tasks vs MoE competitors.
35.7%
43
Google · HuggingFace/Qwen · 2026-04-18
MoE 4B active/26B total. Struggles with complex agentic coding tasks.
13.8%