ExploitBench
2 models tested · Updated 2026-06-09 · Verified sources only
Claude Mythos 5 leads at 78.0%
1
Anthropic · Blog/OpenAI · 2026-06-09
Cyber-only Mythos 5 beats GPT-5.6 Sol (73.5) and Mythos Preview (74.2); reflects lifted cyber safeguards. Fable 5 (with safeguards) scores 40%.
78.0%
2
OpenAI · Blog/OpenAI · 2026-07-09
ExploitBench (V8 exploit chain); 1.5x GPT-5.5 (47.9%) at same token budget.
73.5%