benchmark
.
space
benchmarks
rankings
compare
voices
transcripts
papers
articles
HumanEval+ leaderboard
HumanEval+
2 models tested · Updated 2026-07-14 · Verified sources only
Bonsai 27B (Ternary)
leads at
93.9%
1
Bonsai 27B (Ternary)
PrismML ·
Whitepaper/PrismML
· 2026-07-14
Ternary variant. Coding barely degrades vs FP16 Qwen3.6-27B (95.12) — execution-based pass@1 held through 1.71 bits/weight.
93.9%
2
Bonsai 27B (1-bit)
PrismML ·
Whitepaper/PrismML
· 2026-07-14
1-bit variant. Binary {-1,+1} weights still pass 89.6% of HumanEval+ — coding survives extreme 1-bit quantization.
89.63%