benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
AA-Briefcase leaderboard
AA-Briefcase
11 models tested · Updated 2026-07-21 · Verified sources only
Claude Fable 5
leads at
1583.0Elo
1
Claude Fable 5
Anthropic ·
BenchmarkList/AA
· 2026-07-21
#1 on AA-Briefcase Elo. Long-horizon knowledge work with thousands of fragmented files.
1583.0Elo
2
Kimi K3
Moonshot ·
BenchmarkList/AA
· 2026-07-21
#2 on AA-Briefcase Elo.
1543.0Elo
3
GPT-5.6 Sol
OpenAI ·
BenchmarkList/AA
· 2026-07-21
#3 on AA-Briefcase Elo.
1496.0Elo
4
Claude Sonnet 5
Anthropic ·
BenchmarkList/AA
· 2026-07-21
#4 on AA-Briefcase Elo.
1388.0Elo
5
Claude Opus 4.8
Anthropic ·
BenchmarkList/AA
· 2026-07-21
#5 on AA-Briefcase Elo.
1354.0Elo
6
Grok 4.5
xAI ·
BenchmarkList/AA
· 2026-07-21
#6 on AA-Briefcase Elo.
1323.0Elo
7
GLM-5.2
Z AI ·
Artificial Analysis
· 2026-07-21
Long-horizon knowledge work benchmark. 454 Elo above Qwen 3.6 27B.
1260.0Elo
8
GPT-5.5
OpenAI ·
BenchmarkList/AA
· 2026-07-21
#7 on AA-Briefcase Elo.
1154.0Elo
9
MiniMax M3
MiniMax ·
BenchmarkList/AA
· 2026-07-21
#8 on AA-Briefcase Elo.
1110.0Elo
10
DeepSeek V4 Pro
DeepSeek ·
BenchmarkList/AA
· 2026-07-21
#9 on AA-Briefcase Elo.
932.0Elo
11
Qwen 3.6 27B
Alibaba ·
Artificial Analysis
· 2026-07-21
Long-horizon knowledge work benchmark. 454 Elo below GLM-5.2.
806.0Elo