benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
ARC-AGI 3 leaderboard
ARC-AGI 3
8 models tested · Updated 2026-07-09 · Verified sources only
GPT-5.6 Sol
leads at
7.78%
1
GPT-5.6 Sol
OpenAI ·
Blog/OpenAI
· 2026-07-09
ARC-AGI-3 SOTA; ~5x the next-best model (Opus 4.8 at 1.5%).
7.78%
2
GPT-5.6 Terra
OpenAI ·
Blog/OpenAI
· 2026-07-09
Terra scores 0.8% on ARC-AGI-3 vs Sol's 7.78% — abstract reasoning is Sol-only strength.
0.8%
3
Claude Opus 4.6
Anthropic ·
ARC Prize/arcprize.org
· 2026-03-24
Score updated after ARC Prize changed scoring normalization (median player baseline, 115% cap per level). Was 0.2% under old scoring.
0.5%
4
Gemini 3.1 Pro Preview
Google ·
ARC Prize/arcprize.org
· 2026-03-24
Highest score on new interactive reasoning benchmark. Humans score 100%. ARC-AGI-3 uses turn-based game environments with no instructions.
0.37%
5
Gemini 3.1 Pro
Google ·
arxiv/2603.24621
· 2026-03-24
All frontier models below 1% on ARC-AGI-3. Interactive reasoning requiring exploration and continuous learning. Humans solve 100%.
0.37%
6
GPT-5.4
OpenAI ·
ARC Prize/arcprize.org
· 2026-03-24
Second-highest on ARC-AGI-3. All frontier models score below 1% on this new interactive reasoning benchmark.
0.26%
7
GPT-5.6 Luna
OpenAI ·
Blog/OpenAI
· 2026-07-09
Luna scores 0.18% on ARC-AGI-3; abstract reasoning reserved for Sol tier.
0.18%
8
Grok 4.20
xAI ·
ARC Prize Foundation
· 2026-03-25
Only frontier model to score exactly 0% on ARC-AGI 3 at launch. Exceeded action cutoff on every level. Humans score 100%.
0.0%