BrowseComp
45 models tested · Updated 2026-07-09 · Verified sources only
GPT-5.6 Sol Ultra leads at 92.2%
1
OpenAI · Blog/OpenAI · 2026-07-09
Ultra mode (4 parallel agents) sets BrowseComp SOTA at 92.2%.
92.2%
2
Moonshot AI · Blog/Moonshot AI · 2026-07-16
Context compaction at 300K tokens. 90.4 with no context mgmt.
91.2%
3
OpenAI · Blog/OpenAI · 2026-07-09
Beats Mythos 5 (88.0%) and Fable 5 (84.3%); Ultra mode hits 92.2%.
90.4%
4
Moonshot AI · Blog/Moonshot AI · 2026-07-16
First open 3T-class model (2.8T params). Score achieved with 1M-token context window and no context management (vs Claude-style compaction).
90.4%
5
OpenAI · OpenAI Blog · 2026-04-23
New BrowseComp SOTA. Nearly 5pt over Gemini 3.1 Pro.
90.1%
6
OpenAI · Blog/OpenAI · 2026-03-05
Highest BrowseComp score reported. Pro variant with maximum reasoning.
89.3%
7
Anthropic · Blog/OpenAI · 2026-06-09
Above Mythos Preview (87.9) and Fable 5 (84.3); below GPT-5.6 Sol (90.4) and Sol Ultra (92.2).
88.0%
8
Anthropic · Blog/Anthropic · 2026-06-09
Score as cited in Kimi K3 blog (source: Anthropic official). Fable 5 hit fallbacks on 35% of tasks.
88.0%
9
OpenAI · Blog/OpenAI · 2026-07-09
Terra surpasses GPT-5.5 (84.4%) on BrowseComp at lower cost.
87.5%
10
Anthropic · Blog/Anthropic · 2026-04-07
New SOTA. Uses 4.9x fewer tokens than Opus 4.6 while scoring higher.
86.9%
11
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
Agent Swarm mode with parallel sub-agents. 2nd highest BrowseComp score after GPT-5.4 Pro.
86.3%
12
Google · Blog/Google · 2026-02-19
Up from 59.2% on Gemini 3 Pro. Strong autonomous web research capability.
85.9%
13
Google · Blog/Google · 2026-04-21
New BrowseComp SOTA for autonomous research. Uses Gemini 3.1 Pro with extended search and reasoning.
85.9%
14
OpenAI · OpenAI Blog · 2026-04-23
Beats GPT-5.4 (82.7) but trails Gemini 3.1 Pro (85.9) and GPT-5.5 Pro (90.1).
84.4%
15
Anthropic · Blog/Anthropic · 2026-06-17
Score as cited in Kimi K3 blog (source: Anthropic official).
84.3%
16
Multi-agent web search coordination across hours-long sessions. 86.8% with harness.
84.0%
17
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning. Competitive with GPT-5.4 (82.7), below GPT-5.5 Pro (90.1).
83.4%
18
OpenAI · Blog/OpenAI · 2026-07-09
Luna nearly matches GPT-5.5 peak (84.4%) at less than half the cost.
83.3%
19
Moonshot AI · HuggingFace/moonshotai · 2026-04-20
Open-weight 1T/32B MoE. Strong web research score, beating GPT-5.4 at 82.7%.
83.2%
20
OpenAI · Blog/OpenAI · 2026-03-05
Up from 65.8% on GPT-5.2. Native computer-use model with 1M context.
82.7%
21
Zhipu AI · HuggingFace/zai-org · 2026-04-07
With context management. No-tools version: 68.0%. Strong browsing capabilities.
79.3%
22
Anthropic · Blog/Anthropic · 2026-04-16
Notable regression from Opus 4.6 (83.7%). First major BrowseComp decline in a frontier model update.
79.3%
23
Alibaba · HuggingFace/Qwen · 2026-02-16
With aggressive context-folding strategy. Outperforms every US frontier model on web browsing.
78.6%
24
Moonshot AI · Blog/Kimi · 2026-01-27
With Agent Swarm. Base score 60.6, with context management 74.9. Competitive with GPT-5.4 (82.7).
78.4%
25
Moonshot AI · Paper/Moonshot AI · 2026-02-01
Multi-agent swarm configuration. 72.7% on WideSearch.
78.4%
26
Thinking Machines Lab · Blog/Thinking Machines Lab · 2026-07-20
Strong browsing for open-weights; below GPT-5.6 Sol Ultra (92.2) but above Kimi K2.5 (74.9).
77.1%
27
MiniMax · Blog/MiniMax · 2026-02-12
Strong browsing agent performance. Competitive with Claude Opus 4.6 (84.0%).
76.3%
28
Zhipu AI · Paper/Zhipu AI (arxiv:2602.15763) · 2026-02-11
With context management strategy. Baseline without CM: 62.0%. Highest open-model BrowseComp score reported.
75.9%
29
StepFun · Blog/StepFun · 2026-05-29
BrowseComp. Approaches Claude Opus 4.7 (79.3) at Flash scale.
75.8%
30
Anthropic · Blog/Anthropic · 2026-02-17
Single-agent config. Corrected after cheating detection pipeline update (was 74.72%, adjusted to 74.01%). Multi-agent variant scores 82.07%.
74.0%
31
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning, 13B activated. Respectable for small model.
73.2%
32
StepFun · Blog/StepFun · 2026-02-12
Score with Context Manager agent framework. Base score 51.6. Strong browsing for an open model.
69.0%
33
Zhipu AI · Blog/Z.AI · 2026-04-07
Top open-model score on BrowseComp. 79.3 with context management variant.
68.0%
34
Alibaba · HuggingFace/Qwen · 2026-02-16
Small model (27B params) with competitive browsing ability. Below frontier but strong for open-weight.
61.0%
35
Moonshot AI · Paper/Moonshot AI · 2026-02-01
Single-agent mode. Agent Swarm achieves 78.4%.
60.6%
36
arxiv · arxiv/2603.15594 · 2026-03-16
First fully open-source search agent achieving frontier-level performance. OpenSeeker-30B trained on only 11.7k samples achieves 29.5% on BrowseComp and 48.4% on BrowseComp-ZH, surpassing industrial c
54.9%
37
Sarvam AI · HuggingFace/sarvamai · 2026-03-06
India's first domestically-trained 105B model. MoE with 10.3B active params. Apache 2.0.
49.5%
38
arxiv · arxiv/2603.15594 · 2026-03-16
First fully open-source search agent achieving frontier-level performance. OpenSeeker-30B trained on only 11.7k samples achieves 29.5% on BrowseComp and 48.4% on BrowseComp-ZH, surpassing industrial c
49.1%
39
Xiaomi · HuggingFace/XiaomiMiMo-MiMo-V2-Flash · 2026-01-06
BrowseComp baseline (no context management).
45.4%
40
Zhipu AI · HuggingFace/Zhipu · 2026-01-15
Lightweight flash variant browser capability score.
42.8%
41
arxiv · arxiv/2603.15594 · 2026-03-16
First fully open-source search agent achieving frontier-level performance. OpenSeeker-30B trained on only 11.7k samples achieves 29.5% on BrowseComp and 48.4% on BrowseComp-ZH, surpassing industrial c
35.3%
42
arxiv · arxiv/2603.15594 · 2026-03-16
First fully open-source search agent achieving frontier-level performance. OpenSeeker-30B trained on only 11.7k samples achieves 29.5% on BrowseComp and 48.4% on BrowseComp-ZH, surpassing industrial c
29.5%
43
arxiv · arxiv/2603.28376 · 2026-03-30
8B deep research agent with verification-centric design surpasses 30B-scale agents on BrowseComp. RL stage adds +6.7 on xBench-DeepSearch.
17.3%
44
arxiv · arxiv/2603.28376 · 2026-03-30
8B deep research agent with verification-centric design surpasses 30B-scale agents on BrowseComp. RL stage adds +6.7 on xBench-DeepSearch.
16.5%
45
arxiv · arxiv/2603.15594 · 2026-03-16
First fully open-source search agent achieving frontier-level performance. OpenSeeker-30B trained on only 11.7k samples achieves 29.5% on BrowseComp and 48.4% on BrowseComp-ZH, surpassing industrial c
15.3%