benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
ExploitBench leaderboard
ExploitBench
2 models tested · Updated 2026-06-09 · Verified sources only
Claude Mythos 5
leads at
78.0%
1
Claude Mythos 5
Anthropic ·
Blog/OpenAI
· 2026-06-09
Cyber-only Mythos 5 beats GPT-5.6 Sol (73.5) and Mythos Preview (74.2); reflects lifted cyber safeguards. Fable 5 (with safeguards) scores 40%.
78.0%
2
GPT-5.6 Sol
OpenAI ·
Blog/OpenAI
· 2026-07-09
ExploitBench (V8 exploit chain); 1.5x GPT-5.5 (47.9%) at same token budget.
73.5%