Terminal-Bench 2.1
27 models tested · Updated 2026-07-09 · Verified sources only
GPT-5.6 Sol Ultra leads at 91.9%
1
OpenAI · Blog/OpenAI · 2026-07-09
Ultra mode (4 parallel agents) sets Terminal-Bench 2.1 SOTA.
91.9%
2
OpenAI · Blog/OpenAI · 2026-07-09
Edges Claude Mythos 5 (88.0%); Ultra mode hits 91.9%.
88.8%
3
Moonshot AI · Blog/Moonshot AI · 2026-07-16
KimiCode harness.
88.3%
4
Anthropic · Blog/Anthropic · 2026-06-09
Safeguard-affected score (~Mythos 5). Terminal-Bench 2.1, not 2.0.
88.0%
5
Anthropic · Blog/OpenAI · 2026-06-09
Matches GPT-5.6 Sol (88.8) on Terminal-Bench 2.1; Mythos 5 is the unsafeguarded twin of Fable 5 (83.1).
88.0%
6
OpenAI · Blog/OpenAI · 2026-07-09
Terra beats GPT-5.5 (85.6%) on Terminal-Bench 2.1.
87.4%
7
Anthropic · Blog/Z.ai · 2026-06-13
Beaten only by Fable 5 (88.0). Top open/commercial.
85.0%
8
OpenAI · Blog/OpenAI · 2026-07-09
Luna near-parity with GPT-5.5 (85.6%) at lowest cost.
84.7%
9
Anthropic · Blog/Anthropic · 2026-06-09
Score as cited in Kimi K3 blog (source: Anthropic official). Fable 5 hit fallbacks on 35% of tasks.
84.6%
10
Anthropic · Blog/Anthropic · 2026-06-17
Score as cited in Kimi K3 blog (source: Anthropic official).
84.6%
11
OpenAI · Blog/OpenAI · 2026-06-26
Score as cited in Kimi K3 blog (source: OpenAI official).
83.4%
12
xAI · Blog/xAI · 2026-07-08
Within 1 point of Fable 5 (max) and GPT-5.5 (xhigh) at 84.3/83.4. First xAI model trained specifically for coding/agent tasks, 500K context.
83.3%
13
Z.ai · Blog/Z.ai · 2026-06-13
Score as cited in Kimi K3 blog (source: Z.ai official).
82.7%
14
Z.ai · Blog/Z.ai · 2026-06-13
Strongest open-source model. Terminus-2 harness. Up from GLM-5.1 (63.5).
81.0%
15
Google DeepMind · Blog/Google DeepMind · 2026-06-10
Terminus-2 harness. Below GPT-5.5 (78.2). Up from Gemini 3 Flash (58.0).
76.2%
16
Alibaba · Blog/Z.ai · 2026-06-13
Below GLM-5.2 (81.0) and Gemini 3.1 Pro (74).
75.0%
17
Moonshot AI · Blog/ThinkingMachines · 2026-07-15
From TML Inkling blog comparison table.
71.3%
18
Google DeepMind · Blog/Google DeepMind · 2026-06-10
Terminus-2. Below Gemini 3.5 Flash (76.2).
70.3%
19
Anthropic · Blog/Google DeepMind · 2026-06-10
Terminus-2. Lower than Gemini 3.5 Flash (76.2).
66.1%
20
MiniMax · Blog/MiniMax · 2026-06-01
Official score 66.0 (Terminus-2). Internal infra, 8C16G sandbox, 2hr timeout.
66.0%
21
MiniMax · Blog/Z.ai · 2026-06-13
Below GLM-5.2 (81.0). Open-weight 1M context model.
65.0%
22
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Uses internal coding harness. Below GLM-5.2 (82.7) but above Nemotron 3 Ultra (56.4).
63.8%
23
DeepSeek · Blog/StepFun-Step-3.7-Flash · 2026-05-29
Terminal-Bench 2.1 score from StepFun internal testing.
62.0%
24
StepFun · Blog/StepFun · 2026-05-29
Terminal-Bench 2.1. +6.1 over Step 3.5 Flash.
59.6%
25
Google DeepMind · Blog/Google DeepMind · 2026-06-10
Terminus-2. Below Gemini 3.5 Flash (76.2).
58.0%
26
NVIDIA · Blog/ThinkingMachines · 2026-07-15
From TML Inkling blog.
56.4%
27
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
Lower than Inkling (63.8) on Terminal-Bench.
52.7%