benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
RealWorldQA leaderboard
RealWorldQA
9 models tested · Updated 2026-04-07 · Verified sources only
Qwen 3.5 9B
leads at
85.7%
1
Qwen 3.5 9B
Alibaba ·
HF/Qwen
· 2026-04-07
Top score among compared small models. Beats Qwen3-VL-30B-A3B (81.9) and Gemini-2.5-Flash-Lite (78.9).
85.7%
2
Qwen 3.6 Plus
Alibaba ·
Blog/Alibaba
· 2026-03-31
Leading on real-world image reasoning, ahead of Gemini 3 Pro (83.3).
85.4%
3
Qwen 3.6 35B-A3B
Alibaba ·
HuggingFace/Qwen
· 2026-04-16
Best among compared models. Beats Claude Sonnet 4.5 (70.3) and Gemma 4 31B (72.3).
85.3%
4
Qwen 3.6 27B
Alibaba ·
HuggingFace/Qwen
· 2026-04-22
Strong real-world visual QA, beats Gemma 4 31B by 12pp.
84.1%
5
Qwen 3.5 397B
Alibaba ·
HuggingFace/Qwen
· 2026-02-16
Real-world visual understanding. Trail Qwen3.5 9B (85.7) — smaller model may have targeted optimization for this benchmark.
83.9%
6
STEP3-VL 10B
StepFun ·
arxiv/2601.09668
· 2026-01-14
10B open-source model outperforms many larger VLMs on real-world visual reasoning.
76.93%
7
Penguin-VL 8B
Apple ·
arxiv/2603.06569
· 2026-03-06
8B model exceeds Qwen3-VL 8B (71.5) and InternVL3.5 8B (67.5) on RealWorldQA using LLM-based vision encoder.
75.8%
8
Vero Q3I-8B
Vero Team ·
arxiv/2604.04917
· 2026-04-06
+1.8 over base. Real-world visual question answering.
73.3%
9
Vero Mi-7B
Vero Team ·
arxiv/2604.04917
· 2026-04-06
+3.0 over MiMoVL base. Real-world visual understanding improvement via RL.
70.6%