benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
SimpleQA Verified leaderboard
SimpleQA Verified
5 models tested · Updated 2026-04-08 · Verified sources only
Muse Spark
leads at
66.0%
1
Muse Spark
Meta ·
X/@EpochAIResearch
· 2026-04-08
Independent evaluation by Epoch AI. Factual accuracy benchmark.
66.0%
2
DeepSeek V4 Pro
DeepSeek ·
DeepSeek/HuggingFace
· 2026-04-24
Open-source leader. 2nd only to Gemini 3.1 Pro (75.6).
57.9%
3
Inkling
Thinking Machines Lab ·
Blog/Thinking Machines Lab
· 2026-07-20
Factuality; below GPT-5.6 Sol (71.6) but above Nemotron 3 Ultra (32.4). Calibration-trained.
43.9%
4
DeepSeek V4 Flash
DeepSeek ·
HuggingFace/deepseek-ai
· 2026-04-24
Factual accuracy for 13B activated model.
34.1%
5
Inkling-Small
Thinking Machines Lab ·
Blog/ThinkingMachines
· 2026-07-15
Much lower factuality than Inkling (43.9).
20.9%