benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
ExploitGym leaderboard
ExploitGym
1 models tested · Updated 2026-07-09 · Verified sources only
GPT-5.6 Sol
leads at
33.7%
1
GPT-5.6 Sol
OpenAI ·
Blog/OpenAI
· 2026-07-09
ExploitGym 6-hour cap; 2.2x GPT-5.5 (15.1%). 24.9% at 2-hour cap.
33.7%