SimpleQA Verified
5 models tested · Updated 2026-04-08 · Verified sources only
Muse Spark leads at 66.0%
1
Meta · X/@EpochAIResearch · 2026-04-08
Independent evaluation by Epoch AI. Factual accuracy benchmark.
66.0%
2
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Open-source leader. 2nd only to Gemini 3.1 Pro (75.6).
57.9%
3
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Factuality; below GPT-5.6 Sol (71.6) but above Nemotron 3 Ultra (32.4). Calibration-trained.
43.9%
4
DeepSeek · HuggingFace/deepseek-ai · 2026-04-24
Factual accuracy for 13B activated model.
34.1%
5
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
Much lower factuality than Inkling (43.9).
20.9%