benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
head to head
Model comparisons
190 matchups across 125 models. Click any to see the full breakdown.
GPT-5.5
OpenAI
11
-
17
28 benchmarks
Claude Opus 4.8
Anthropic
GPT-5.5
OpenAI
2
-
26
28 benchmarks
GPT-5.6 Sol
OpenAI
Claude Opus 4.8
Anthropic
7
-
19
27 benchmarks
GPT-5.6 Sol
OpenAI
GPT-5.5
OpenAI
1
-
24
26 benchmarks
Claude Fable 5
Anthropic
Claude Opus 4.8
Anthropic
1
-
24
26 benchmarks
Claude Fable 5
Anthropic
GPT-5.6 Sol
OpenAI
10
-
16
26 benchmarks
Claude Fable 5
Anthropic
GPT-5.5
OpenAI
1
-
23
25 benchmarks
Kimi K3
Moonshot AI
Claude Opus 4.8
Anthropic
3
-
22
25 benchmarks
Kimi K3
Moonshot AI
GPT-5.6 Sol
OpenAI
9
-
15
25 benchmarks
Kimi K3
Moonshot AI
Claude Fable 5
Anthropic
14
-
10
25 benchmarks
Kimi K3
Moonshot AI
Claude Opus 4.8
Anthropic
17
-
4
21 benchmarks
GLM-5.2
Z.ai
GPT-5.5
OpenAI
13
-
5
18 benchmarks
GLM-5.2
Z.ai
Claude Opus 4.7
Anthropic
15
-
3
18 benchmarks
Claude Opus 4.6
Anthropic
GPT-5.6 Sol
OpenAI
17
-
0
17 benchmarks
GLM-5.2
Z.ai
Gemini 3.1 Pro
Google
8
-
9
17 benchmarks
GPT-5.4
OpenAI
Claude Fable 5
Anthropic
17
-
0
17 benchmarks
GLM-5.2
Z.ai
Gemini 3.1 Pro
Google
8
-
8
16 benchmarks
Claude Opus 4.7
Anthropic
Kimi K3
Moonshot AI
16
-
0
16 benchmarks
GLM-5.2
Z.ai
GPT-5.5
OpenAI
11
-
4
15 benchmarks
Gemini 3.1 Pro
Google
GPT-5.5
OpenAI
12
-
3
15 benchmarks
Claude Opus 4.7
Anthropic
Gemma 4 31B
Google
15
-
0
15 benchmarks
Gemma 4 26B A4B
Google
GPT-5.5
OpenAI
13
-
0
14 benchmarks
GPT-5.4
OpenAI
Gemini 3.1 Pro
Google
8
-
5
14 benchmarks
Claude Opus 4.6
Anthropic
GPT-5.4
OpenAI
7
-
7
14 benchmarks
Claude Opus 4.6
Anthropic
Claude Opus 4.7
Anthropic
8
-
5
13 benchmarks
GPT-5.4
OpenAI
Claude Opus 4.7
Anthropic
8
-
4
12 benchmarks
Kimi K2.5
Moonshot AI
Claude Opus 4.7
Anthropic
9
-
3
12 benchmarks
Kimi K2.6
Moonshot AI
Qwen 3.6 27B
Alibaba
11
-
1
12 benchmarks
Qwen 3.6 35B-A3B
Alibaba
GPT-5.6 Sol
OpenAI
10
-
1
11 benchmarks
Gemini 3.1 Pro
Google
Gemini 3.1 Pro
Google
10
-
1
11 benchmarks
Kimi K2.5
Moonshot AI
Gemini 3.1 Pro
Google
8
-
3
11 benchmarks
Kimi K2.6
Moonshot AI
Claude Opus 4.7
Anthropic
9
-
2
11 benchmarks
Gemma 4 31B
Google
Claude Opus 4.7
Anthropic
7
-
4
11 benchmarks
DeepSeek V4 Pro
DeepSeek
Claude Opus 4.7
Anthropic
11
-
0
11 benchmarks
Gemma 4 26B A4B
Google
Claude Opus 4.6
Anthropic
9
-
2
11 benchmarks
Kimi K2.5
Moonshot AI
Claude Opus 4.6
Anthropic
9
-
2
11 benchmarks
Gemma 4 31B
Google
Claude Opus 4.6
Anthropic
9
-
2
11 benchmarks
Gemma 4 26B A4B
Google
Kimi K2.5
Moonshot AI
9
-
2
11 benchmarks
Qwen 3.6 35B-A3B
Alibaba
GPT-5.5
OpenAI
8
-
1
10 benchmarks
Claude Opus 4.6
Anthropic
GPT-5.5
OpenAI
9
-
0
10 benchmarks
Kimi K2.6
Moonshot AI
Claude Opus 4.8
Anthropic
4
-
6
10 benchmarks
Gemini 3.1 Pro
Google
Claude Opus 4.8
Anthropic
9
-
1
10 benchmarks
Inkling
Thinking Machines Lab
GPT-5.6 Sol
OpenAI
8
-
2
10 benchmarks
Claude Opus 4.7
Anthropic
Gemini 3.1 Pro
Google
1
-
9
10 benchmarks
Claude Fable 5
Anthropic
Gemini 3.1 Pro
Google
10
-
0
10 benchmarks
Qwen3.5 27B
Alibaba
Gemini 3.1 Pro
Google
10
-
0
10 benchmarks
Gemma 4 31B
Google
Gemini 3.1 Pro
Google
10
-
0
10 benchmarks
Gemma 4 26B A4B
Google
Gemini 3.1 Pro
Google
9
-
1
10 benchmarks
Inkling
Thinking Machines Lab
Claude Opus 4.7
Anthropic
9
-
1
10 benchmarks
Qwen 3.6 35B-A3B
Alibaba
Claude Opus 4.6
Anthropic
6
-
4
10 benchmarks
DeepSeek V4 Pro
DeepSeek
Claude Opus 4.6
Anthropic
6
-
4
10 benchmarks
Kimi K2.6
Moonshot AI
Claude Opus 4.6
Anthropic
7
-
3
10 benchmarks
Qwen 3.6 35B-A3B
Alibaba
Kimi K2.5
Moonshot AI
9
-
1
10 benchmarks
Qwen3.5 27B
Alibaba
Kimi K2.5
Moonshot AI
10
-
0
10 benchmarks
Gemma 4 31B
Google
Kimi K2.5
Moonshot AI
0
-
10
10 benchmarks
Kimi K2.6
Moonshot AI
Kimi K2.5
Moonshot AI
10
-
0
10 benchmarks
Gemma 4 26B A4B
Google
Gemma 4 31B
Google
1
-
8
10 benchmarks
Qwen 3.6 35B-A3B
Alibaba
Kimi K2.6
Moonshot AI
10
-
0
10 benchmarks
Qwen 3.6 27B
Alibaba
Gemma 4 26B A4B
Google
0
-
10
10 benchmarks
Qwen 3.6 35B-A3B
Alibaba
GPT-5.5
OpenAI
8
-
1
9 benchmarks
DeepSeek V4 Pro
DeepSeek
GPT-5.5
OpenAI
9
-
0
9 benchmarks
Inkling
Thinking Machines Lab
Claude Opus 4.8
Anthropic
6
-
3
9 benchmarks
Claude Opus 4.7
Anthropic
Claude Opus 4.8
Anthropic
7
-
2
9 benchmarks
Kimi K2.6
Moonshot AI
GPT-5.6 Sol
OpenAI
9
-
0
9 benchmarks
Inkling
Thinking Machines Lab
Gemini 3.1 Pro
Google
7
-
1
9 benchmarks
DeepSeek V4 Pro
DeepSeek
Gemini 3.1 Pro
Google
9
-
0
9 benchmarks
Qwen 3.6 27B
Alibaba
Gemini 3.1 Pro
Google
9
-
0
9 benchmarks
Qwen 3.6 35B-A3B
Alibaba
Claude Fable 5
Anthropic
8
-
1
9 benchmarks
Claude Opus 4.7
Anthropic
Claude Fable 5
Anthropic
9
-
0
9 benchmarks
Inkling
Thinking Machines Lab
Claude Opus 4.7
Anthropic
9
-
0
9 benchmarks
Inkling
Thinking Machines Lab
Claude Opus 4.7
Anthropic
8
-
1
9 benchmarks
Qwen 3.6 27B
Alibaba
GPT-5.4
OpenAI
9
-
0
9 benchmarks
Kimi K2.5
Moonshot AI
Claude Opus 4.6
Anthropic
6
-
3
9 benchmarks
Qwen 3.6 27B
Alibaba
Kimi K2.5
Moonshot AI
0
-
9
9 benchmarks
DeepSeek V4 Pro
DeepSeek
Kimi K2.5
Moonshot AI
5
-
4
9 benchmarks
Qwen 3.6 27B
Alibaba
Gemma 4 31B
Google
0
-
9
9 benchmarks
DeepSeek V4 Pro
DeepSeek
Gemma 4 31B
Google
1
-
8
9 benchmarks
Qwen 3.6 27B
Alibaba
DeepSeek V4 Pro
DeepSeek
9
-
0
9 benchmarks
Gemma 4 26B A4B
Google
Kimi K2.6
Moonshot AI
8
-
1
9 benchmarks
Inkling
Thinking Machines Lab
Kimi K2.6
Moonshot AI
9
-
0
9 benchmarks
Qwen 3.6 35B-A3B
Alibaba
Gemma 4 26B A4B
Google
0
-
9
9 benchmarks
Qwen 3.6 27B
Alibaba
GPT-5.5
OpenAI
8
-
0
8 benchmarks
Kimi K2.5
Moonshot AI
GPT-5.5
OpenAI
8
-
0
8 benchmarks
Gemma 4 31B
Google
GPT-5.5
OpenAI
8
-
0
8 benchmarks
Gemma 4 26B A4B
Google
Claude Opus 4.8
Anthropic
8
-
0
8 benchmarks
DeepSeek V4 Pro
DeepSeek
GPT-5.6 Sol
OpenAI
8
-
0
8 benchmarks
GPT-5.4
OpenAI
GPT-5.6 Sol
OpenAI
8
-
0
8 benchmarks
Kimi K2.6
Moonshot AI
Gemini 3.1 Pro
Google
2
-
6
8 benchmarks
Kimi K3
Moonshot AI
Gemini 3.1 Pro
Google
4
-
4
8 benchmarks
GLM-5.2
Z.ai
Kimi K3
Moonshot AI
8
-
0
8 benchmarks
Inkling
Thinking Machines Lab
Claude Opus 4.7
Anthropic
8
-
0
8 benchmarks
Qwen3.5 27B
Alibaba
GLM-5.2
Z.ai
4
-
4
8 benchmarks
DeepSeek V4 Pro
DeepSeek
GLM-5.2
Z.ai
8
-
0
8 benchmarks
Inkling
Thinking Machines Lab
GPT-5.4
OpenAI
5
-
3
8 benchmarks
Kimi K2.6
Moonshot AI
Claude Opus 4.6
Anthropic
6
-
2
8 benchmarks
Qwen3.5 27B
Alibaba
EXAONE 4.5 33B
LG AI Research
2
-
6
8 benchmarks
Qwen3.5 27B
Alibaba
EXAONE 4.5 33B
LG AI Research
1
-
7
8 benchmarks
Qwen 3.6 35B-A3B
Alibaba
Gemma 4 31B
Google
0
-
8
8 benchmarks
Kimi K2.6
Moonshot AI
DeepSeek V4 Pro
DeepSeek
8
-
0
8 benchmarks
Qwen 3.6 35B-A3B
Alibaba
Kimi K2.6
Moonshot AI
8
-
0
8 benchmarks
Gemma 4 26B A4B
Google