benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
research
Papers
81 papers with benchmark results from arxiv, tagged by theme
Small Beats Big
11
Agent
12
Coding
11
OCR
10
Efficient
5
Frontier
5
Reasoning
4
VLM
3
Browser Agent
2
Distillation
2
Small VLM
2
Test Time Compute
2
Autoregressive
1
Data Curation
1
Diffusion LLM
1
Model Release
1
MOE
1
OCR Document
1
Quantization
1
Rlvr
1
Safety Evaluation
1
Small Models
1
Tool Use
1
Vision Language
1
Filtering by:
Coding
·
clear
Coding
11 papers
ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision?
2026-04-09 · arxiv
Coding
Small Model
Rlvr
Awakening the Sleeping Agent: Lean-Specific Agentic Data Reactivates General Tool Use in Goedel Prover
2026-04-09 · arxiv
Coding
Formal Math
Model Release
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
2026-04-09 · arxiv
Coding
Swe Bench
CODESTRUCT: Code Agents over Structured Action Spaces
2026-04-07 · arxiv
Coding
InCoder-32B-Thinking: Industrial Code World Model for Thinking
2026-04-03 · arxiv
Coding
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
2026-04-02 · arxiv
Coding
Yet Even Less Is Even Better For Agentic, Reasoning, and Coding LLMs
2026-04-01 · arxiv
Coding
KAT-Coder-V2 Technical Report
2026-03-29 · arxiv
Coding
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code
2026-03-06 · arxiv
Coding
Qwen3-Coder-Next Technical Report
2026-02-28 · arxiv
Coding
SWE-Protege: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents
2026-02-25 · arxiv
Coding