RealWorldQA
9 models tested · Updated 2026-04-07 · Verified sources only
Qwen 3.5 9B leads at 85.7%
1
Alibaba · HF/Qwen · 2026-04-07
Top score among compared small models. Beats Qwen3-VL-30B-A3B (81.9) and Gemini-2.5-Flash-Lite (78.9).
85.7%
2
Alibaba · Blog/Alibaba · 2026-03-31
Leading on real-world image reasoning, ahead of Gemini 3 Pro (83.3).
85.4%
3
Alibaba · HuggingFace/Qwen · 2026-04-16
Best among compared models. Beats Claude Sonnet 4.5 (70.3) and Gemma 4 31B (72.3).
85.3%
4
Alibaba · HuggingFace/Qwen · 2026-04-22
Strong real-world visual QA, beats Gemma 4 31B by 12pp.
84.1%
5
Alibaba · HuggingFace/Qwen · 2026-02-16
Real-world visual understanding. Trail Qwen3.5 9B (85.7) — smaller model may have targeted optimization for this benchmark.
83.9%
6
StepFun · arxiv/2601.09668 · 2026-01-14
10B open-source model outperforms many larger VLMs on real-world visual reasoning.
76.93%
7
Apple · arxiv/2603.06569 · 2026-03-06
8B model exceeds Qwen3-VL 8B (71.5) and InternVL3.5 8B (67.5) on RealWorldQA using LLM-based vision encoder.
75.8%
8
Vero Team · arxiv/2604.04917 · 2026-04-06
+1.8 over base. Real-world visual question answering.
73.3%
9
Vero Team · arxiv/2604.04917 · 2026-04-06
+3.0 over MiMoVL base. Real-world visual understanding improvement via RL.
70.6%