research
Papers
81 papers with benchmark results from arxiv, tagged by theme
Small Beats Big 11Agent 12Coding 11OCR 10Efficient 5Frontier 5Reasoning 4VLM 3Browser Agent 2Distillation 2Small VLM 2Test Time Compute 2Autoregressive 1Data Curation 1Diffusion LLM 1Model Release 1MOE 1OCR Document 1Quantization 1Rlvr 1Safety Evaluation 1Small Models 1Tool Use 1Vision Language 1
Filtering by: Agent · clear
Agent
12 papers
Structured Distillation of Web Agent Capabilities Enables Generalization
2026-04-09 · arxiv
Agent
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
2026-04-06 · arxiv
AgentTool Use
Beyond Fixed Tests: Repository-Level Issue Resolution as Coevolution of Code and Behavioral Constraints
2026-04-06 · arxiv
AgentCoding
InfoSeeker: A Scalable Hierarchical Parallel Agent Framework for Web Information Seeking
2026-04-03 · arxiv
AgentSearch AgentMulti Agent
Marco DeepResearch: Unlocking Efficient Deep Research Agents via Verification-Centric Design
2026-03-30 · arxiv
AgentSearch AgentVerification
Seed1.8 Model Card: Towards Generalized Real-World Agency
2026-03-21 · arxiv
AgentFoundation ModelMultimodal
WebNavigator: Global Web Navigation via Interaction Graph Retrieval
2026-03-20 · arxiv
AgentWeb AgentNavigation
A Subgoal-driven Framework for Improving Long-Horizon LLM Agents
2026-03-20 · arxiv
AgentWeb AgentRlSmall Model
OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data
2026-03-16 · arxiv
AgentSearch AgentOpen Source
Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents (GUI-Owl-1.5)
2026-02-15 · arxiv
Agent
OpAgent: Operator Agent for Web Navigation
2026-02-14 · arxiv
Agent
Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use Agents
2026-02-03 · arxiv
Agent