WebArena
18 models tested · Updated 2026-02-14 · Verified sources only
OpAgent leads at 71.6%
1
Research · arxiv/2602.13559 · 2026-02-14
New WebArena SOTA. Modular planner/grounder/reflector/summarizer with online RL. Previous SOTA was ColorBrowserAgent at 71.2%.
71.6%
2
H Company · Blog/H Company · 2025-10-17
Agent architecture with separated planning and execution. Higher than raw model scores (GPT-5.4 at 67.3) due to agentic orchestration.
69.6%
3
OpenAI · Blog/OpenAI · 2026-03-05
WebArena-Verified. Uses both DOM- and screenshot-driven interaction. Improves over GPT-5.2 (65.4%).
67.3%
4
Moonshot AI · Blog/Kimi · 2026-01-27
Web browsing agent benchmark.
58.9%
5
arxiv · arxiv/2603.20366 · 2026-03-20
WebNavigator reframes web navigation as deterministic retrieval over Interaction Graphs, achieving 72.9% on WebArena multi-site tasks and 39.7% on OnlineMind2Web, more than doubling enterprise-level a
57.1%
6
Google DeepMind · arxiv/2604.07776 · 2026-04-09
From arxiv paper on structured distillation. Tested with Gemini 3.1 Pro API as web agent.
53.8%
7
Google DeepMind · arxiv/2604.07776 · 2026-04-09
From arxiv paper on structured distillation. Gemini 3 Pro as web agent.
51.2%
8
arxiv · arxiv/2603.20366 · 2026-03-20
WebNavigator reframes web navigation as deterministic retrieval over Interaction Graphs, achieving 72.9% on WebArena multi-site tasks and 39.7% on OnlineMind2Web, more than doubling enterprise-level a
49.9%
9
Alibaba · arxiv/2602.16855 · 2026-02-15
Best open-weight WebArena result for 32B class. Thinking variant for web navigation.
48.4%
10
arxiv · arxiv/2603.20366 · 2026-03-20
WebNavigator reframes web navigation as deterministic retrieval over Interaction Graphs, achieving 72.9% on WebArena multi-site tasks and 39.7% on OnlineMind2Web, more than doubling enterprise-level a
47.8%
11
Alibaba · arxiv/2602.16855 · 2026-02-15
Best open-source result on WebArena. Beats proprietary models by wide margin for its size.
46.7%
12
McGill-NLP · arxiv/2604.07776 · 2026-04-09
Distilled from Gemini 3 Pro using structured trajectory synthesis. Matches 3x larger Qwen3.5-27B and nearly doubles previous best open-weight SFT result (21.7%).
41.5%
13
Alibaba · arxiv/2604.07776 · 2026-04-09
Teacher model for A3 distillation. Matched by 9B student on WebArena.
41.5%
14
arxiv · arxiv/2604.07776 · 2026-04-09
Agent-as-Annotators framework distills web agent capabilities from Gemini 3 Pro into a 9B-param student. A3-Qwen3.5-9B matches Qwen3.5-27B on WebArena (41.5% SR), beating GPT-4o (31.5%).
41.5%
15
arxiv · arxiv/2604.07776 · 2026-04-09
Agent-as-Annotators framework distills web agent capabilities from Gemini 3 Pro into a 9B-param student. A3-Qwen3.5-9B matches Qwen3.5-27B on WebArena (41.5% SR), beating GPT-4o (31.5%).
36.0%
16
A3 Team · arxiv/2604.07776 · 2026-04-09
4B distilled model beats GPT-4o (31.5%). +11.1pp over base Qwen3.5-4B via structured distillation.
35.2%
17
arxiv · arxiv/2604.07776 · 2026-04-09
Agent-as-Annotators framework distills web agent capabilities from Gemini 3 Pro into a 9B-param student. A3-Qwen3.5-9B matches Qwen3.5-27B on WebArena (41.5% SR), beating GPT-4o (31.5%).
31.5%
18
arxiv · arxiv/2604.07776 · 2026-04-09
Agent-as-Annotators framework distills web agent capabilities from Gemini 3 Pro into a 9B-param student. A3-Qwen3.5-9B matches Qwen3.5-27B on WebArena (41.5% SR), beating GPT-4o (31.5%).
31.0%