benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
research
Papers
81 papers with benchmark results from arxiv, tagged by theme
Small Beats Big
11
Agent
12
Coding
11
OCR
10
Efficient
5
Frontier
5
Reasoning
4
VLM
3
Browser Agent
2
Distillation
2
Small VLM
2
Test Time Compute
2
Autoregressive
1
Data Curation
1
Diffusion LLM
1
Model Release
1
MOE
1
OCR Document
1
Quantization
1
Rlvr
1
Safety Evaluation
1
Small Models
1
Tool Use
1
Vision Language
1
Filtering by:
VLM
·
clear
VLM
3 papers
Vero: An Open RL Recipe for General Visual Reasoning
2026-04-06 · arxiv
VLM
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
2026-03-06 · arxiv
VLM
STEP3-VL-10B Technical Report
2026-01-14 · arxiv
VLM