benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
SWE-bench Pro leaderboard
SWE-bench Pro
43 models tested · Updated 2026-06-09 · Verified sources only
Claude Fable 5
leads at
80.3%
1
Claude Fable 5
Anthropic ·
Blog/Anthropic
· 2026-06-09
New SOTA. 21.7 points above GPT-5.5 (58.6%). Agentic coding at scale.
80.3%
2
Claude Mythos 5
Anthropic ·
Blog/OpenAI
· 2026-06-09
Unsafeguarded Mythos-class model via Project Glasswing. Same underlying model as Fable 5 with cyber safeguards lifted.
80.3%
3
Claude Mythos Preview
Anthropic ·
Blog/Anthropic
· 2026-04-07
24.4pp above Opus 4.6 (53.4%). Massive jump on harder coding eval.
77.8%
4
Claude Opus 4.8
Anthropic ·
Blog/Z.ai
· 2026-06-13
Second only to Fable 5 (80.3). Beats GLM-5.2 (62.1).
69.2%
5
Grok 4.5
xAI ·
Blog/xAI
· 2026-07-08
Third on SWE-bench Pro behind Fable 5 (80.4) and Opus 4.8 (69.2). Notable for token efficiency: 4.2x fewer output tokens than Opus 4.8.
64.7%
6
GPT-5.6 Sol
OpenAI ·
Blog/OpenAI
· 2026-07-09
Beats GPT-5.5 (59.4%); trails Claude Mythos 5 (80.3%).
64.6%
7
Claude Opus 4.7
Anthropic ·
Blog/Anthropic
· 2026-04-16
Up 10.9pp from Opus 4.6. Largest single-version jump on SWE-bench Pro.
64.3%
8
GPT-5.6 Terra
OpenAI ·
Blog/OpenAI
· 2026-07-09
Balanced-tier Terra beats GPT-5.5 (59.4%) at lower cost.
63.4%
9
GPT-5.6 Luna
OpenAI ·
Blog/OpenAI
· 2026-07-09
Cheapest tier Luna beats GPT-5.5 (59.4%) on SWE-bench Pro.
62.7%
10
GLM-5.2
Z.ai ·
Blog/Z.ai
· 2026-06-13
Strongest open-source score. Closes gap to Opus 4.8 (69.2). Up from GLM-5.1 (58.4).
62.1%
11
Qwen3.7-Max
Alibaba ·
Blog/Z.ai
· 2026-06-13
Above GPT-5.5 (58.6), below GLM-5.2 (62.1).
60.6%
12
MiniMax M3
MiniMax ·
Blog/Z.ai
· 2026-06-13
Above GPT-5.5 (58.6). Competitive with GLM-5.2 (62.1).
59.0%
13
Kimi K2.6
Moonshot AI ·
HuggingFace/moonshotai
· 2026-04-20
Open-weight SOTA on SWE-bench Pro. 1T/32B MoE matching GPT-5.4 at 57.7%.
58.6%
14
GPT-5.5
OpenAI ·
OpenAI Blog
· 2026-04-23
Matches Kimi K2.6 SOTA. Uses fewer tokens than GPT-5.4.
58.6%
15
GPT-5.5 Pro
OpenAI ·
OpenAI Blog
· 2026-04-23
Same as GPT-5.5, listed separately in Pro column.
58.6%
16
GLM-5.1
Z.ai ·
X/@Zai_org
· 2026-04-07
#1 open source, #3 globally. Beats GPT-5.4 (57.7) and Claude Opus 4.6 (57.3). First Chinese model to top SWE-bench Pro. Zero Nvidia hardware.
58.4%
17
Hunyuan Hy3
Tencent ·
HuggingFace/tencent-Hy3-modelcard
· 2026-07-06
SWE-bench Pro. Competitive with DeepSeek V4 Pro (55.4).
57.9%
18
GPT-5.4
OpenAI ·
Blog/OpenAI
· 2026-03-05
Harder coding benchmark variant. GPT-5.4 splits coding, reasoning, and computer use into one unified model.
57.7%
19
Claude Opus 4.6
Anthropic ·
Z.ai/GLM-5.1-blog
· 2026-04-08
Cited alongside GPT-5.4 (57.7) and GLM-5.1 (58.4) in multiple independent comparisons. Opus 4.6 was previous SOTA before GLM-5.1.
57.3%
20
Qwen 3.6 Max Preview
Alibaba ·
X/@Alibaba_Qwen
· 2026-04-20
Barely edges past Opus 4.5 (57.1). Top coding model from China. Proprietary, not open-weight.
57.3%
21
GPT-5.3 Codex
OpenAI ·
Web/OpenAI
· 2026-03-05
Leads SWE-bench Pro. 25% faster than GPT-5.2 Codex. Covers 1,865 multi-language tasks.
56.8%
22
Step 3.7 Flash
StepFun ·
Blog/StepFun
· 2026-05-29
SWE-bench Pro. 2nd place among open models, behind DeepSeek V4 Pro (55.4) is actually ahead.
56.3%
23
MiniMax M2.7
MiniMax ·
Blog/MiniMax
· 2026-03-18
Matches GPT-5.3 Codex level. Nearly approaches Opus best on SWE-Pro.
56.2%
24
DeepSeek V4 Pro
DeepSeek ·
DeepSeek/HuggingFace
· 2026-04-24
Max reasoning. Below GPT-5.5 (58.6) and Kimi K2.6 (58.6).
55.4%
25
Gemini 3.5 Flash
Google DeepMind ·
Blog/Google DeepMind
· 2026-06-10
Public set. Below Opus 4.7 (64.3). Up from Gemini 3 Flash (49.6).
55.1%
26
GPT-5.4 Mini
OpenAI ·
Blog/OpenAI
· 2026-03-17
Smallest GPT-5.4 variant. 2x faster than GPT-5 mini at similar coding quality.
54.4%
27
Inkling
Thinking Machines Lab ·
Blog/Thinking Machines Lab
· 2026-07-20
Open-weights MoE; trails GLM-5.2 (62.1) and Fable 5 (80.0) on harder SWE variant.
54.3%
28
Gemini 3.1 Pro
Google ·
Blog/Google DeepMind
· 2026-02-19
Competitive with GPT-5.4 (57.7%). Behind Claude Mythos Preview (77.8%).
54.2%
29
Qwen 3.6 27B
Alibaba ·
HuggingFace/Qwen
· 2026-04-22
Competitive with Qwen 3.5 397B MoE on agent coding.
53.5%
30
Claude Opus 4.6
Anthropic ·
arxiv/Mythos-System-Card
· 2026-04-07
Scored by Anthropic using Scale AI variant. Lower than GPT-5.4 (57.7%) and Gemini 3.1 Pro (54.2%) on this harder variant.
53.4%
31
Inkling-Small
Thinking Machines Lab ·
Blog/ThinkingMachines
· 2026-07-15
SWE-bench Pro Public.
53.2%
32
DeepSeek V4 Flash
DeepSeek ·
HuggingFace/deepseek-ai
· 2026-04-24
Impressive for 13B activated — competitive with much larger models.
52.6%
33
GPT-5.4 Nano
OpenAI ·
Blog/OpenAI
· 2026-03-17
Fastest, cheapest GPT-5.4 variant. Designed for classification, data extraction, and coding subagents.
52.4%
34
Muse Spark
Meta ·
Blog/Meta
· 2026-04-08
Trails Claude Mythos (77.8%) significantly but competitive with other frontier models.
52.0%
35
Kimi K2.5
Moonshot AI ·
HuggingFace/Moonshot
· 2026-01-29
Internal evaluation framework with minimal tool set. 1T param MoE, 32B active.
50.7%
36
Qwen 3.6 35B-A3B
Alibaba ·
HuggingFace/Qwen
· 2026-04-16
3B active params MoE beating Gemma 4 31B (35.7%) on SWE-bench Pro.
49.5%
37
Seed 2.0 Pro
ByteDance ·
Blog/ByteDance
· 2026-02-14
Flagship coding model from ByteDance Seed team.
46.9%
38
Nemotron 3 Ultra
NVIDIA ·
Blog/ThinkingMachines
· 2026-07-15
From TML Inkling blog. SWE-bench Pro Public.
46.4%
39
Seed 2.0 Lite
ByteDance ·
Blog/ByteDance
· 2026-03-10
Nearly matches Pro variant (46.9) on harder SWE subset.
46.0%
40
Qwen 3 Coder Next 80B A3B
Alibaba ·
Blog/Alibaba
· 2026-02-03
SWE-Agent score. Competitive with much larger models on harder SWE subset.
44.3%
41
Qwen 3 Coder Next 80B A3B
Qwen ·
arxiv/2603.00729
· 2026-02-28
3B active params matching models with 10x more active compute on SWE-bench Pro.
42.7%
42
Gemma 4 31B
Google ·
HuggingFace/Qwen
· 2026-04-02
Struggles on harder SWE tasks vs MoE competitors.
35.7%
43
Gemma 4 26B A4B
Google ·
HuggingFace/Qwen
· 2026-04-18
MoE 4B active/26B total. Struggles with complex agentic coding tasks.
13.8%