MedXPertQA MM
4 models tested · Updated 2026-04-02 · Verified sources only
Gemma 4 31B leads at 61.3%
1
Google · HuggingFace/Google DeepMind · 2026-04-02
Medical multimodal QA benchmark.
61.3%
2
Google · HuggingFace/Google DeepMind · 2026-04-02
MoE 25.2B total, 3.8B active.
58.1%
3
Google DeepMind · HuggingFace/google-gemma-4-12b-it · 2026-05-23
Medical expert QA multimodal.
48.7%
4
LG AI Research · HuggingFace/LGAI-EXAONE · 2026-04-14
Beats GPT-5 mini (34.4). Comparable to Qwen3-VL 32B (41.6).
42.1%