benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
OSWorld-Verified leaderboard
OSWorld-Verified
6 models tested · Updated 2026-06-09 · Verified sources only
Claude Fable 5
leads at
85.0%
1
Claude Fable 5
Anthropic ·
Blog/Anthropic
· 2026-06-09
Computer-use SOTA. Beats GPT-5.5 (78.7%) and Opus 4.8 (83.4%).
85.0%
2
Gemini 3.5 Flash
Google DeepMind ·
Blog/Google DeepMind
· 2026-06-10
Agentic computer use. Close to GPT-5.5 (78.7). Up from Gemini 3 Flash (65.1).
78.4%
3
Claude Opus 4.7
Anthropic ·
Blog/Google DeepMind
· 2026-06-10
Agentic computer use. Close to GPT-5.5 (78.7).
78.0%
4
Gemini 3.1 Pro
Google DeepMind ·
Blog/Google DeepMind
· 2026-06-10
Below Gemini 3.5 Flash (78.4).
76.2%
5
MiniMax M3
MiniMax ·
Blog/MiniMax
· 2026-06-01
361 samples (nogdrive). Relative coords 0-1000, 1920x1080, Max Steps 200.
70.06%
6
Gemini 3 Flash
Google DeepMind ·
Blog/Google DeepMind
· 2026-06-10
Below Gemini 3.5 Flash (78.4). 13-point gain.
65.1%