Humanity's Last Exam
72 models tested · Updated 2026-04-07 · Verified sources only
Claude Mythos Preview leads at 64.7%
1
Anthropic · Blog/Anthropic · 2026-04-07
With tools variant. No-tools score is 56.8%. 11.6 pts ahead of Opus 4.6 with tools (53.1%).
64.7%
2
Anthropic · Blog/Anthropic · 2026-06-09
Safeguard-affected (~Mythos 5). GPT-5.5 scores 41.4% on standard config.
59.0%
3
OpenAI · OpenAI Blog · 2026-04-23
With tools. Still HLE-with-tools SOTA, above GPT-5.5 Pro (57.2).
58.7%
4
Alibaba · Blog/Qwen · 2026-01-25
Highest self-reported HLE score. With tools. Leads GPT-5.2-Thinking by 28%. First Chinese model to top HLE.
58.3%
5
Meta · Blog/Meta AI · 2026-04-08
Contemplating mode (multi-agent parallel reasoning). First model from Meta Superintelligence Labs. Accepts voice, text, image inputs. Strong multimodal + health, weaker on agentic tasks vs Mythos/GPT-5.4.
58.0%
6
Anthropic · Blog/Z.ai · 2026-06-13
With tools, full set.
57.9%
7
OpenAI · OpenAI Blog · 2026-04-23
With tools. Second only to GPT-5.4 Pro (58.7).
57.2%
8
Z.ai · Blog/Z.ai · 2026-06-13
With tools, full set. Up from GLM-5.1 (52.3).
54.7%
9
Google · Blog/Google · 2026-04-21
HLE score with tool use (Deep Research Max). Uses extended search/reasoning time.
54.6%
10
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
New open-weight SOTA on HLE with tools. Exceeds GPT-5.4 (52.1%) and Opus 4.6 (53.0%).
54.0%
11
Alibaba · Blog/Z.ai · 2026-06-13
With tools, full set.
53.5%
12
Anthropic · Blog/Anthropic · 2026-06-09
Score as cited in Kimi K3 blog (source: Anthropic official). Fable 5 hit fallbacks on 35% of tasks.
53.3%
13
Tencent · HuggingFace/tencent-Hy3-modelcard · 2026-07-06
HLE score. Strong reasoning for a 21B-active MoE.
53.2%
14
Hardest reasoning benchmark. With tool use. Corrected score per Feb 23 update.
53.1%
15
OpenAI · OpenAI Blog · 2026-04-23
With tools. Near GPT-5.4 Pro (58.7) and Opus 4.6 (54.7).
52.2%
16
OpenAI · arxiv/Mythos-System-Card · 2026-04-07
With tools, from Anthropic system card.
52.1%
17
Google · arxiv/Mythos-System-Card · 2026-04-07
With tools from Anthropic system card.
51.4%
18
Alibaba · Blog/Qwen · 2026-04-02
Strong HLE score, competitive with frontier models on this extreme reasoning test.
50.6%
19
Zhipu AI · HuggingFace/zai-org · 2026-02-11
Tool-augmented evaluation variant. Base (no tools) score is 30.5%. Higher than Claude Opus 4.5 tools score.
50.4%
20
Anthropic · Blog/Z.ai · 2026-06-13
Leads non-Mythos models on HLE. Close to Fable 5 safeguarded score.
49.8%
21
Alibaba · HuggingFace/Qwen · 2026-02-16
With tool access. Nearly 2x the CoT-only score. Impressive for a 27B model — beats several frontier models.
48.5%
22
Google · Blog/Google · 2026-02-12
Without tools. Google specialized reasoning mode. Also scored 84.6% on ARC-AGI-2.
48.4%
23
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
With tools, Max reasoning. Below leaders but competitive.
48.2%
24
StepFun · Blog/StepFun · 2026-05-29
HLE with tools. Flash-tier model beating Step 3.5 (35.7) and DeepSeek V4 Flash (45.1).
47.2%
25
Anthropic · Blog/Google DeepMind · 2026-06-10
Full set text+MM. Highest among non-Mythos models.
46.9%
26
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
HLE with tools. Inkling-Small preview (276B MoE, 12B active) matches its larger sibling on reasoning.
46.6%
27
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Open-weights 975B MoE (41B active). First model from Thinking Machines Lab (Mira Murati). Multimodal from scratch.
46.0%
28
OpenAI · Blog/OpenAI · 2026-07-09
Score as cited in Kimi K3 blog (source: OpenAI official).
44.5%
29
Google · Google DeepMind — Gemini page · 2026-02-19
Without tools. Google DeepMind official benchmarks page. Leads frontier models on HLE standard eval.
44.4%
30
OpenAI · Scale AI Leaderboard · 2026-03-05
Top of Scale AI HLE leaderboard. Standard evaluation, no tools.
44.32%
31
Moonshot AI · Blog/Moonshot AI · 2026-07-16
HLE-Full (no tools).
43.5%
32
OpenAI · OpenAI Blog · 2026-04-23
No tools. Still below Opus 4.6 (46.9) and Gemini 3.1 Pro (44.4).
43.1%
33
Zhipu AI · NVIDIA/Zhipu Official · 2025-12-22
Open-weight model with tools, competitive with GPT-5.1 High (42.7%).
42.8%
34
OpenAI · OpenAI Blog · 2026-04-23
No tools. Below GPT-5.5 Pro (43.1).
42.7%
35
OpenAI · OpenAI Blog · 2026-04-23
No tools. Below Opus 4.6 (46.9) and Gemini 3.1 Pro (44.4).
41.4%
36
Alibaba · Blog/Z.ai · 2026-06-13
Ties GPT-5.5 (41.4). Text-only subset.
41.4%
37
Z.ai · Blog/Z.ai · 2026-06-13
Text-only subset. Up from GLM-5.1 (31). 10-point jump.
40.5%
38
Google DeepMind · Blog/Google DeepMind · 2026-06-10
Full set text+MM. Below Opus 4.7 (46.9). Up from Gemini 3 Flash (33.7).
40.2%
39
Anthropic · arxiv/Mythos-System-Card · 2026-04-07
No tools. Below GPT-5.4 (39.8%) and Gemini 3.1 Pro (44.4%).
40.0%
40
OpenAI · Leaderboard/Artificial Analysis · 2026-03-24
Second-highest on HLE behind GPT-5.4 Pro (44.3%). Codex-native agent combining frontier coding with general reasoning.
39.9%
41
OpenAI · arxiv/Mythos-System-Card · 2026-04-07
No tools, from Anthropic system card. Between Claude Opus (40.0%) and Gemini (44.4%).
39.8%
42
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
No tools, Max reasoning. Below GPT-5.5 (41.4) and Gemini 3.1 Pro (44.4).
37.7%
43
Google · Scale AI Leaderboard · 2025-11-18
Second on Scale AI HLE leaderboard. Standard evaluation, no tools.
37.52%
44
MiniMax · Blog/Z.ai · 2026-06-13
Text-only. Below GLM-5.2 (40.5) and Qwen3.7-Max (41.4).
37.0%
45
OpenAI · Scale AI Leaderboard · 2026-03-05
xhigh thinking config on Scale AI HLE leaderboard.
36.24%
46
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
No tools, Max reasoning, 13B activated. Near Kimi K2.6 (34.7).
34.8%
47
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
HLE without tools. Improved from K2.5 (30.1%) but below GPT-5.4 (39.8%) and Opus 4.6 (40.0%).
34.7%
48
Google · Blog/Google · 2025-12-17
Frontier intelligence at Flash speed. Outperforms Gemini 2.5 Pro on HLE without tools.
33.7%
49
Anthropic · Blog/Google DeepMind · 2026-06-10
Full set text+MM. Lowest among compared frontier models.
33.2%
50
OpenAI · Scale AI Leaderboard · 2025-10-06
Standard evaluation on Scale AI HLE leaderboard.
31.64%
51
Zhipu AI · HuggingFace/zai-org · 2026-04-07
No-tools. Slight improvement over GLM-5 (30.5%). 754B open-weight MoE.
31.0%
52
Zhipu AI · Paper/Zhipu AI (arxiv:2602.15763) · 2026-02-11
Text-only subset. With tools: 50.4% (comparable to Kimi K2.5 51.8%).
30.5%
53
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
HLE text-only. Inkling 975B MoE (41B active), Thinking Machines Lab first open-weights release.
29.7%
54
Thinking Machines Lab · Blog/ThinkingMachines · 2026-07-15
HLE text-only variant.
29.6%
55
OpenAI · Scale AI Leaderboard · 2025-12-11
Standard evaluation on Scale AI HLE leaderboard.
27.8%
56
Google · HuggingFace/Google DeepMind · 2026-04-02
With search tools enabled. Without search: 19.5%.
26.5%
57
Moonshot · Scale AI Leaderboard · 2026-01-27
Chinese frontier model. Strong HLE showing for open-weight model.
24.37%
58
Alibaba · HuggingFace/Qwen · 2026-02-16
Chain-of-thought only. With tools jumps to 48.5%. Shows HLE requires tool use for small models.
24.3%
59
xAI · Artificial Analysis · 2025-07-09
Frontier reasoning benchmark. Third-party verified via Artificial Analysis.
24.0%
60
Alibaba · HuggingFace/Qwen · 2026-04-22
Solid HLE for a small dense model.
24.0%
61
Xiaomi · HuggingFace/XiaomiMiMo-MiMo-V2-Flash · 2026-01-06
HLE no tools. Low but non-zero for Flash-tier.
22.1%
62
Alibaba · HuggingFace/Qwen · 2026-04-16
Open MoE with competitive HLE score. Beats Qwen 3.5 35B-A3B (22.4) is slightly lower.
21.4%
63
arxiv · arxiv/2604.06753 · 2026-04-08
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
20.0%
64
Google · HuggingFace/Google · 2026-04-02
Without tools. Dense 31B model. Strong HLE for an open-weight model.
19.5%
65
MiniMax · Blog/MiniMax · 2026-02-12
HLE without tools. Behind frontier models (Opus 4.6: 30.7, GPT-5.2: 31.4).
19.4%
66
NVIDIA · arxiv/2603.19220 · 2026-03-19
HLE no-tool score; lower than Qwen3.5-35B (22.4), showing frontier reasoning still difficult for 3B active params.
17.7%
67
Google · HuggingFace/Google DeepMind · 2026-04-02
With search tools enabled. Without search: 8.7%.
17.2%
68
Zhipu AI · HuggingFace/Zhipu · 2026-01-15
Without tools. 30B-A3B MoE, 3.6B active params.
14.4%
69
arxiv · arxiv/2604.06753 · 2026-04-08
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
14.0%
70
arxiv · arxiv/2604.06753 · 2026-04-08
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
14.0%
71
arxiv · arxiv/2604.06753 · 2026-04-08
Independent evaluation of 4 frontier LLMs across 6 reasoning paradigms and 10 benchmarks. No single paradigm dominates; learned router recovers 37% of oracle gap.
14.0%
72
Google DeepMind · HuggingFace/google-gemma-4-12b-it · 2026-05-23
HLE no tools. Low but non-zero for 12B.
5.2%