benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
research
Papers
81 papers with benchmark results from arxiv, tagged by theme
Small Beats Big
11
Agent
12
Coding
11
OCR
10
Efficient
5
Frontier
5
Reasoning
4
VLM
3
Browser Agent
2
Distillation
2
Small VLM
2
Test Time Compute
2
Autoregressive
1
Data Curation
1
Diffusion LLM
1
Model Release
1
MOE
1
OCR Document
1
Quantization
1
Rlvr
1
Safety Evaluation
1
Small Models
1
Tool Use
1
Vision Language
1
Filtering by:
Frontier
·
clear
Frontier
5 papers
Dead Weights, Live Signals: Feedforward Graphs of Frozen Language Models
2026-04-09 · arxiv
Frontier
Model Combination
Small Model
Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents
2026-04-08 · arxiv
Frontier
Agent
Reasoning
Multi-objective Evolutionary Merging Enables Efficient Reasoning Models
2026-04-07 · arxiv
Frontier
Reasoning
Model Merging
PRISM-MCTS: Learning from Reasoning Trajectories with Metacognitive Reflection
2026-04-06 · arxiv
Frontier
Reasoning
Search
Phi-4-reasoning-vision-15B Technical Report
2026-03-04 · arxiv
Frontier
New Model
VLM
Small Model