Identifies Seeing but Not Thinking phenomenon in multimodal MoE models: vision tokens routed to reasoning experts cause distraction. Soft routing intervention improves Qwen3-VL-30B-A3B average from 59
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
Identifies Seeing but Not Thinking phenomenon in multimodal MoE models: vision tokens routed to reasoning experts cause distraction. Soft routing intervention improves Qwen3-VL-30B-A3B average from 59
Introduces Gaussian GRPO (G2RPO), replacing standard linear scaling in GRPO with non-linear distributional matching. OpenVLThinkerV2-7B achieves new SOTA for open-source 7B models on MMMU (71.6%), Mat
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
Compact 15B open-weight multimodal reasoning model from Microsoft. Achieves competitive VLM performance with much less compute via careful data curation and dynamic-resolution encoders.
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
Proposes RLSD combining self-distillation magnitude with RLVR direction. On Qwen3-VL-8B, achieves best avg accuracy across 5 multimodal reasoning benchmarks, outperforming GRPO by 2.32% on average.
4B end-to-end OCR model that ranks #1 on OmniDocBench among end-to-end models. Introduces "Layout-as-Thought" for structured layout representations. Outperforms Qwen3-VL-4B on ChartQA (+4.8) and Chart
Compact 15B open-weight multimodal reasoning model from Microsoft. Achieves competitive VLM performance with much less compute via careful data curation and dynamic-resolution encoders.
8B open-weight multimodal model trained with GRPO+GDPO. Competitive with Gemini 2.5 Pro on DocVQA and chart understanding. New SOTA for open-weight VLMs on MMMU.
4B end-to-end OCR model that ranks #1 on OmniDocBench among end-to-end models. Introduces "Layout-as-Thought" for structured layout representations. Outperforms Qwen3-VL-4B on ChartQA (+4.8) and Chart
Identifies Seeing but Not Thinking phenomenon in multimodal MoE models: vision tokens routed to reasoning experts cause distraction. Soft routing intervention improves Qwen3-VL-30B-A3B average from 59
Compact 15B open-weight multimodal reasoning model from Microsoft. Achieves competitive VLM performance with much less compute via careful data curation and dynamic-resolution encoders.
Compact 15B open-weight multimodal reasoning model from Microsoft. Achieves competitive VLM performance with much less compute via careful data curation and dynamic-resolution encoders.
Compact 15B open-weight multimodal reasoning model from Microsoft. Achieves competitive VLM performance with much less compute via careful data curation and dynamic-resolution encoders.
0.8B parameter VLM using Mamba-based SSM with cross-modality modulation instead of cross-attention. Outperforms MobileVLM V2 (1.7B) and MoE-LLaVA (2.2B) on several benchmarks despite being much smalle