benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
research
Papers
81 papers with benchmark results from arxiv, tagged by theme
Small Beats Big
11
Agent
12
Coding
11
OCR
10
Efficient
5
Frontier
5
Reasoning
4
VLM
3
Browser Agent
2
Distillation
2
Small VLM
2
Test Time Compute
2
Autoregressive
1
Data Curation
1
Diffusion LLM
1
Model Release
1
MOE
1
OCR Document
1
Quantization
1
Rlvr
1
Safety Evaluation
1
Small Models
1
Tool Use
1
Vision Language
1
Filtering by:
Agent
·
clear
Agent
12 papers
Structured Distillation of Web Agent Capabilities Enables Generalization
2026-04-09 · arxiv
Agent
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
2026-04-06 · arxiv
Agent
Tool Use
Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints
2026-04-06 · arxiv
Agent
Coding
InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking
2026-04-03 · arxiv
Agent
Search Agent
Multi Agent
Marco DeepResearch: Unlocking Efficient Deep Research Agents via Verification-Centric Design
2026-03-30 · arxiv
Agent
Search Agent
Verification
Seed1.8 Model Card: Towards Generalized Real-World Agency
2026-03-21 · arxiv
Agent
Foundation Model
Multimodal
WebNavigator: Global Web Navigation via Interaction Graph Retrieval
2026-03-20 · arxiv
Agent
Web Agent
Navigation
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
2026-03-20 · arxiv
Agent
Web Agent
Rl
Small Model
OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data
2026-03-16 · arxiv
Agent
Search Agent
Open Source
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents (GUI-Owl-1.5)
2026-02-15 · arxiv
Agent
OpAgent: Operator Agent for Web Navigation
2026-02-14 · arxiv
Agent
Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use Agents
2026-02-03 · arxiv
Agent