benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
tau2-bench leaderboard
tau2-bench
2 models tested · Updated 2026-03-02 · Verified sources only
Qwen3.5 4B
leads at
79.9%
1
Qwen3.5 4B
Qwen ·
HuggingFace/Qwen
· 2026-03-02
Edges out the 9B on tau2-bench agentic tool-use.
79.9%
2
Qwen 3.5 9B
Alibaba ·
HuggingFace/Qwen
· 2026-03-02
Strong agentic tool-use; near-top among open sub-10B models.
79.1%