OmniDocBench
31 models tested · Updated 2026-04-06 · Verified sources only
MinerU2.5-Pro leads at 95.69%
1
Shanghai AI Lab · arxiv/2604.04771 · 2026-04-06
SOTA on OmniDocBench v1.6 via pure data engineering — no architecture changes. Expanded training from <10M to 65.5M samples.
95.69%
2
Zhipu AI · arxiv/2603.10910 · 2026-03-11
0.9B model (0.4B encoder + 0.5B decoder) achieves #1 on OmniDocBench v1.5. Outperforms models 260x larger including GPT-5.2 and Gemini-3 Pro.
94.62%
3
Baidu · arxiv/2601.21957 · 2026-01-29
0.9B compact VLM with NaViT dynamic resolution. Best text and formula recognition scores despite minimal size. Also introduces Real5-OmniDocBench for real-world robustness.
94.5%
4
Shanghai AI Lab · arxiv/2603.23885 · 2026-03-25
1B-param end-to-end model achieving new top-tier OmniDocBench score with data-training co-design and structure-aware training.
93.75%
5
Tencent · GitHub/opendatalab · 2026-03-31
Tencent 2.5B specialized VLM for document parsing. Second-highest among specialized models.
93.37%
6
Baidu · arxiv/2603.13398 · 2026-03-11
End-to-end unified OCR model. Wins 6/8 benchmarks in its class. Strong chart reasoning capabilities.
93.12%
7
FireRed AI · arxiv/2603.01840 · 2026-03-02
3-stage training: multi-task pre-alignment, specialized SFT, and GRPO RL with format constraints. Qwen3-VL backbone.
92.94%
8
Ant Group · arxiv/2603.11044 · 2026-03-11
Financial-domain document parser with cross-page consolidation and cell-level referencing. Competitive on OmniDocBench while excelling on FinDocBench.
92.8%
9
Alibaba · GitHub/opendatalab · 2026-03-31
Alibaba 4B document parsing model. Strong table and formula recognition.
92.56%
10
AIDC-AI · GitHub/opendatalab · 2026-03-31
30B MoE (3B active) vision-language model. Top proprietary VLM tier on document parsing.
92.36%
11
Moonshot AI · Blog/Moonshot AI · 2026-07-16
91.1%
12
Tencent · GitHub/opendatalab · 2026-03-31
Tencent Hunyuan 1B OCR model. Excellent formula recognition (95.96 CDM).
90.57%
13
OpenDoc · GitHub/opendatalab · 2026-03-31
Tiny 0.1B OCR model matching much larger models. Remarkable efficiency.
90.57%
14
Nanonets · X/@nanonets · 2026-04-06
Commercial OCR API model. Claims world-leading accuracy with confidence scores and structured extraction endpoints.
90.5%
15
Google · GitHub/opendatalab · 2026-03-31
Google Gemini 3 Flash on OmniDocBench v1.5. Strong generalist VLM.
90.37%
16
DocTron · GitHub/opendatalab · 2026-03-31
4B document parsing model with reinforcement learning approach.
90.2%
17
Google · GitHub/opendatalab · 2026-03-31
Google Gemini 3 Pro on OmniDocBench v1.5.
90.17%
18
Alibaba · HuggingFace/Qwen · 2026-04-16
Top OmniDocBench 1.5 score. Slight improvement over Qwen 3.5 35B-A3B (89.3).
89.9%
19
Anthropic · Blog/Anthropic · 2026-06-09
Score as cited in Kimi K3 blog (source: Anthropic official). Fable 5 hit fallbacks on 35% of tasks.
89.8%
20
OpenAI · Blog/OpenAI · 2026-06-26
Score as cited in Kimi K3 blog (source: OpenAI official).
89.4%
21
Moonshot · GitHub/opendatalab · 2026-03-31
Moonshot Kimi K2.5 1T param model on document parsing.
89.33%
22
DeepSeek · GitHub/opendatalab · 2026-03-31
DeepSeek 3B second-gen OCR model. Improves on original DeepSeek-OCR.
89.17%
23
MonkeyOCR · GitHub/opendatalab · 2026-03-31
3B pro variant of MonkeyOCR with strong text recognition.
88.85%
24
ByteDance · GitHub/opendatalab · 2026-03-31
ByteDance 3B OCR model v2. Balanced performance across document types.
88.71%
25
DocTron · GitHub/opendatalab · 2026-03-31
4B document parsing model. Consistent across text, tables, formulas.
88.55%
26
rednote · GitHub/opendatalab · 2026-03-31
3B OCR model from Xiaohongshu (rednote). Strong formula recognition.
88.41%
27
Anthropic · Blog/Anthropic · 2026-06-17
Score as cited in Kimi K3 blog (source: Anthropic official).
87.9%
28
MonkeyOCR · GitHub/opendatalab · 2026-03-31
Compact 1.2B pro variant. Good accuracy for size class.
86.96%
29
OpenAI · Blog/OpenAI · 2026-07-09
Score as cited in Kimi K3 blog (source: OpenAI official).
85.8%
30
OpenAI · GitHub/opendatalab · 2026-03-31
OpenAI GPT-5.2 on OmniDocBench v1.5. Weaker than specialized OCR models.
85.75%
31
LG AI Research · HuggingFace/LGAI-EXAONE · 2026-04-14
Document understanding focus. Beats GPT-5 mini (77.0). Particular strength in Korean document understanding.
81.2%