InfoVQA
8 models tested · Updated 2026-04-09 · Verified sources only
OpenVLThinkerV2 8B leads at 86.4%
1
Shanghai AI Lab · arxiv/2604.08539 · 2026-04-09
Gaussian GRPO-trained multimodal reasoner achieving strong InfoVQA. Uses distributional matching for inter-task gradient equity across diverse visual tasks.
86.4%
2
arxiv · arxiv/2604.08539 · 2026-04-09
Introduces Gaussian GRPO (G2RPO), replacing standard linear scaling in GRPO with non-linear distributional matching. OpenVLThinkerV2-7B achieves new SOTA for open-source 7B models on MMMU (71.6%), Mat
86.4%
3
arxiv · arxiv/2604.08539 · 2026-04-09
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
84.2%
4
arxiv · arxiv/2604.08539 · 2026-04-09
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
79.0%
5
Q-Mask Team · arxiv/2604.00161 · 2026-03-31
+1.6 over Qwen2.5-VL-3B baseline. Spatial priors help on info-heavy documents.
78.7%
6
Shanghai Jiao Tong University · arxiv/2603.07494 · 2026-03-08
Layout-aware reasoning with Visual-Semantic Chain. Achieves 78.6 on InfoVQA, outperforming Qwen3-VL-8B by 2.9 points.
78.6%
7
Q-Mask Team · arxiv/2604.00161 · 2026-03-31
+2.0 over Qwen3-VL-2B baseline. Small model benefits most from spatial priors.
74.4%
8
arxiv · arxiv/2603.13398 · 2026-03-11
4B end-to-end OCR model that ranks #1 on OmniDocBench among end-to-end models. Introduces "Layout-as-Thought" for structured layout representations. Outperforms Qwen3-VL-4B on ChartQA (+4.8) and Chart
73.76%