DocVQA
15 models tested · Updated 2026-04-09 · Verified sources only
OpenVLThinkerV2 8B leads at 96.7%
1
Research · arxiv/2604.08539 · 2026-04-09
Beats GPT-5 (91.5) and Gemini 2.5 Pro (92.6) on DocVQA with just 8B params.
96.7%
2
arxiv · arxiv/2604.08539 · 2026-04-09
Introduces Gaussian GRPO (G2RPO), replacing standard linear scaling in GRPO with non-linear distributional matching. OpenVLThinkerV2-7B achieves new SOTA for open-source 7B models on MMMU (71.6%), Mat
96.7%
3
Apple · arxiv/2603.06569 · 2026-03-06
8B model matches Qwen3-VL 8B (96.1) and exceeds InternVL3.5 8B (92.3) on DocVQA with custom LLM-based vision encoder.
96.2%
4
arxiv · arxiv/2604.08539 · 2026-04-09
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
95.9%
5
arxiv · arxiv/2604.08539 · 2026-04-09
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
95.3%
6
Apple · arxiv/2603.06569 · 2026-03-06
2B model exceeds InternVL3.5 2B (89.4) on DocVQA; efficient architecture with LLM-based vision encoder.
94.1%
7
Q-Mask Team · arxiv/2604.00161 · 2026-03-31
Competitive with InternVL3.5-8B on document VQA. Query-driven causal masks for text anchoring.
93.5%
8
Q-Mask Team · arxiv/2604.00161 · 2026-03-31
Matches 8B models on DocVQA despite 4x smaller. Causal query-driven mask decoder.
93.3%
9
Shanghai Jiao Tong University · arxiv/2603.07494 · 2026-03-08
Layout-aware reasoning model with Visual-Semantic Chain supervision. Strong generalization across six document benchmarks with progressive GRPO training.
93.2%
10
Baidu · arxiv/2603.13398 · 2026-03-11
Unified end-to-end model. Strong DocVQA performance from OCR-specialized architecture.
92.8%
11
arxiv · arxiv/2604.08539 · 2026-04-09
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
92.6%
12
arxiv · arxiv/2603.13398 · 2026-03-11
4B end-to-end OCR model that ranks #1 on OmniDocBench among end-to-end models. Introduces "Layout-as-Thought" for structured layout representations. Outperforms Qwen3-VL-4B on ChartQA (+4.8) and Chart
91.85%
13
arxiv · arxiv/2604.08539 · 2026-04-09
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
91.5%
14
arxiv · arxiv/2603.03975 · 2026-03-04
Compact 15B open-weight multimodal reasoning model from Microsoft. Achieves competitive VLM performance with much less compute via careful data curation and dynamic-resolution encoders.
89.2%
15
arxiv · arxiv/2603.03975 · 2026-03-04
Compact 15B open-weight multimodal reasoning model from Microsoft. Achieves competitive VLM performance with much less compute via careful data curation and dynamic-resolution encoders.
76.0%