benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
BFCL v3 leaderboard
BFCL v3
2 models tested · Updated 2026-07-14 · Verified sources only
Bonsai 27B (Ternary)
leads at
74.41%
1
Bonsai 27B (Ternary)
PrismML ·
Whitepaper/PrismML
· 2026-07-14
Ternary variant. Tool-calling accuracy — agentic workflows are the design target and hold within a few points of FP16 (77.10).
74.41%
2
Bonsai 27B (1-bit)
PrismML ·
Whitepaper/PrismML
· 2026-07-14
1-bit variant. Tool-calling retained at 70.7 — enables on-device agentic loops on a phone.
70.72%