FrontierMath
11 models tested · Updated 2026-07-09 · Verified sources only
GPT-5.6 Sol leads at 89.0%
1
OpenAI · Blog/OpenAI · 2026-07-09
FrontierMath v2 Tier 1-3, beats Fable 5 (87%) and GPT-5.5 (85.3%).
89.0%
2
OpenAI · Blog/OpenAI · 2026-07-09
FrontierMath v2 Tier 1-3; nearly matches Fable 5 (87%).
84.9%
3
OpenAI · Blog/OpenAI · 2026-07-09
FrontierMath v2 Tier 4 (hardest); trails Fable 5 (87.8%) but nearly doubles GPT-5.5 (72.5%).
83.0%
4
OpenAI · Blog/OpenAI · 2026-07-09
Luna trails GPT-5.5 (85.3%) on FrontierMath v2 Tier 1-3.
78.6%
5
OpenAI · OpenAI Blog · 2026-04-23
FrontierMath Tier 1-3. New SOTA.
52.4%
6
OpenAI · OpenAI Blog · 2026-04-23
FrontierMath Tier 1-3. 4pt over GPT-5.4.
51.7%
7
OpenAI · Epoch AI Blog · 2026-03-05
New FrontierMath record on Tiers 1-3 (undergrad to postdoc math). Also scored 38% on Tier 4 (research-grade). Solved 2 previously unsolved Tier 4 problems.
50.0%
8
OpenAI · OpenAI Blog · 2026-04-23
Tier 1-3. 4pt below GPT-5.5.
47.6%
9
Anthropic · Blog/OpenAI · 2026-04-16
OpenAI-tested. Strong math but below GPT-5.5 and DeepSeek V4 Pro.
43.8%
10
Meta · X/@EpochAIResearch · 2026-04-08
Independent evaluation by Epoch AI on Tiers 1-3 (undergrad to early postdoc). Behind GPT-5.4 Pro (50%) but competitive with other frontier models.
39.0%
11
Google · X/@richa_lq · 2026-04-18
Per @richa_lq (26 likes, 2.3k impressions). OpenAI leads FrontierMath at 50%+ but has access to full dataset.
36.9%