benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
MedXPertQA MM leaderboard
MedXPertQA MM
4 models tested · Updated 2026-04-02 · Verified sources only
Gemma 4 31B
leads at
61.3%
1
Gemma 4 31B
Google ·
HuggingFace/Google DeepMind
· 2026-04-02
Medical multimodal QA benchmark.
61.3%
2
Gemma 4 26B A4B
Google ·
HuggingFace/Google DeepMind
· 2026-04-02
MoE 25.2B total, 3.8B active.
58.1%
3
Gemma 4 12B
Google DeepMind ·
HuggingFace/google-gemma-4-12b-it
· 2026-05-23
Medical expert QA multimodal.
48.7%
4
EXAONE 4.5 33B
LG AI Research ·
HuggingFace/LGAI-EXAONE
· 2026-04-14
Beats GPT-5 mini (34.4). Comparable to Qwen3-VL 32B (41.6).
42.1%