LiveCodeBench
53 models tested · Updated 2026-04-03 · Verified sources only
Arcee Trinity leads at 98.2%
1
Arcee AI · Blog/Arcee AI · 2026-04-03
399B sparse MoE (13B active), Apache 2.0. Impressive score but one user reports 60% failure rate on production constraint-following tasks. Self-reported by Arcee.
98.2%
2
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning. New SOTA, 2pt over Gemini 3.1 Pro (91.7).
93.5%
3
Google · Leaderboard/LiveCodeBench · 2025-11-18
Current #1 on LiveCodeBench leaderboard. Ahead of DeepSeek V3.2 Speciale (88.7).
91.7%
4
DeepSeek · DeepSeek/HuggingFace · 2026-04-24
Max reasoning, 13B activated. Near V4 Pro (93.5) SOTA.
91.6%
5
DeepSeek · Paper/DeepSeek · 2025-12-01
Pass@1 with chain-of-thought. Temporary API-only reasoning variant. Gold-level results in IMO, ICPC World Finals, IOI 2025.
88.7%
6
ByteDance · Blog/ByteDance · 2026-02-14
Leads LiveCodeBench v6 at time of release. ByteDance frontier model with ICPC/IMO/CMO gold medals.
87.8%
7
NVIDIA · arxiv/2603.19220 · 2026-03-19
Strong coding performance from 3B active params; LiveCodeBench v6 (2408-2505). With TIR: 88.4.
87.2%
8
Alibaba · Blog/Qwen · 2026-04-02
Top-4 on LiveCodeBench leaderboard. Strong agentic coding model with 1M context window.
87.1%
9
StepFun · Blog/StepFun · 2026-02-12
Reaches 88.9 with PaCoRe. Open-weight model competitive with frontier closed models on live coding.
86.4%
10
Xiaomi · arxiv/2601.02780 · 2026-01-06
15B active MoE. v6 score. Strong coding performance for size class.
85.1%
11
Moonshot AI · HuggingFace/moonshotai · 2026-01-27
LiveCodeBench v6, 64k token budget. Highest open-source score on this benchmark. Averaged over 5 runs.
85.0%
12
Zhipu AI · NVIDIA/Zhipu Official · 2025-12-22
LiveCodeBench v6. Highest open-weight LiveCodeBench score, beats most proprietary models.
84.9%
13
Alibaba · HuggingFace/Qwen · 2026-02-16
LiveCodeBench v6. Outperforms GPT-5.2 and Claude Opus 4.5 on 80% of evaluated categories.
83.6%
14
DeepSeek · Paper/DeepSeek-V3.2 · 2025-12-01
Pass@1 with chain-of-thought. Open-weight.
83.3%
15
Anthropic · Blog/Anthropic · 2026-04-16
Up 7pp from Opus 4.6 (76.0%). Strong coding competition performance.
83.0%
16
PrismML · Whitepaper/PrismML · 2026-07-14
Ternary variant. Beats Gemma-4-31B Q2_K_XL (74.40) at one-third the footprint — long-form code generation survives ternary compression while conventional sub-4-bit collapses here.
82.75%
17
ByteDance · Blog/ByteDance · 2026-03-10
Strong competitive coding despite being the mid-tier variant.
81.7%
18
LG AI Research · HuggingFace/LGAI-EXAONE · 2026-04-14
First open-weight VLM from LG. Beats Qwen3-VL 235B Thinking (70.1) on LiveCodeBench v6. 33B dense model.
81.4%
19
Kuaishou · arxiv/2604.03144 · 2026-04-03
Highest among open-weight models. +28% over non-thinking variant via Error-driven Chain-of-Thought reasoning.
81.3%
20
Alibaba · Web/Alibaba · 2026-02-16
Dense 27B model, all params active. Matches GPT-5-mini on SWE-bench. Best coding score in Qwen 3.5 medium lineup.
80.7%
21
NVIDIA · Blog/NVIDIA · 2026-03-11
Leading open-weight score on LiveCodeBench. 120B MoE with 12B active params.
80.56%
22
Alibaba · HuggingFace/Qwen · 2026-04-16
LiveCodeBench v6. Matches Gemma 4 31B (80.0) with 10x fewer active params.
80.4%
23
Google · Model Card/Google · 2026-04-02
Official Google model card. 31B dense model on LiveCodeBench v6.
80.0%
24
xAI · Blog/xAI · 2025-07-09
From xAI official blog. No-tools evaluation. LiveCodeBench Jan-May 2025.
79.0%
25
Google · Model Card/Google · 2026-04-02
MoE variant. v6. Dense 31B scores 80.0%.
77.1%
26
PrismML · Whitepaper/PrismML · 2026-07-14
1-bit variant. Still beats Qwen3.6-27B IQ2_XXS (56.40) at 2.5x smaller footprint — the key long-form code-gen metric where sub-4-bit collapses.
76.4%
27
Anthropic · Blog/Anthropic · 2026-02-05
Competitive coding on real-world problems. Behind Gemini 3 Pro Preview (91.7%) and DeepSeek V3.2 Speciale (88.7%).
76.0%
28
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
75.9%
29
DeepSeek · HuggingFace/DeepSeek · 2025-12-01
Standard (non-thinking) mode. Thinking mode reaches 83.3%. Speciale variant hits 88.7%.
74.1%
30
Google · Blog/Google · 2026-03-03
Budget-tier model at $0.25/1M input tokens. Competitive coding despite being Google fastest/cheapest model.
72.0%
31
Sarvam AI · HuggingFace/sarvamai · 2026-03-06
India's first domestically-trained 105B model. MoE with 10.3B active params. Apache 2.0.
71.7%
32
ServiceNow · arxiv/2604.02007 · 2026-04-02
15B open-weight. Matches Nemotron-Cascade while generating less than half the output tokens (7.4K vs 16.0K).
70.8%
33
Alibaba Cloud · arxiv/2601.09088 · 2026-01-15
4B model. v5 score. Competitive with models 8x larger.
69.3%
34
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
68.9%
35
NVIDIA · HuggingFace/NVIDIA · 2025-12-15
v6. Outperforms Qwen3 (66.0%) and GPT-OSS (61.0%) at same size class.
68.3%
36
Alibaba · HuggingFace/Qwen · 2026-03-02
9B params. Competitive coding for a small model, though lags behind larger models (GPT-OSS-120B: 82.7).
65.6%
37
JD.com · X/NoPeople404 · 2026-04-06
48B MoE, 2.7B active params. Surpasses GLM-4.7-Flash-Thinking by 2.4% with 85% fewer tokens. Uses FiberPO RL algorithm.
65.6%
38
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
64.2%
39
Zhipu AI · HuggingFace/Zhipu · 2026-01-15
v6 benchmark. Decent coding for a 3.6B active param model.
64.0%
40
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
59.1%
41
Qwen · arxiv/2603.00729 · 2026-02-28
LiveCodeBench v6; beats Qwen3-Coder-480B-A35B (44.93) and Qwen3-Next (51.79) despite 3B active params.
58.93%
42
Alibaba · HuggingFace/Alibaba · 2026-03-02
Competitive coding ability for 4B params. Open-weight, Apache 2.0.
55.8%
43
Research · arxiv/2603.05863 · 2026-03-06
8B open-source model. pass@5 score. Beats GPT-5.1 (48.03) and Claude Sonnet 4.5 (50.78) on this benchmark.
54.12%
44
Google · Google Model Card · 2026-04-02
Competitive coding for a 4B model. Runs on consumer edge hardware.
52.0%
45
arxiv · arxiv/2604.07725 · 2026-04-09
Multi-model orchestration for verifier-free evolutionary inference. RSA baselines show GPT-5 mini at 85.0% GPQA Diamond and 94.2% AIME. Qwen3-235B at 84.3% GPQA Diamond.
46.1%
46
Google · Model Card/Google · 2026-04-02
v6. Effective 2B params. Surpasses Gemma 3 27B on coding despite 10x fewer active parameters.
44.0%
47
Meta · HuggingFace/Meta · 2026-04-05
pass@1 on LiveCodeBench (10/01/2024-02/01/2025). Far behind frontier models like Gemma 4 31B (80%) on coding.
43.4%
48
arxiv · arxiv/2604.01193 · 2026-04-01
Simple self-distillation (sample + fine-tune) improves Qwen3-30B from 42.4% to 55.3% on LiveCodeBench v6. Works across 4B, 8B, and 30B scales. No verifier, teacher model, or RL needed.
42.4%
49
Meta · HuggingFace/Meta · 2026-04-05
pass@1. Significantly behind frontier coding models.
32.8%
50
arxiv · arxiv/2604.07864 · 2026-04-09
RLVR code training without ground-truth tests. ZeroCoder+DyB4 on Qwen2.5-Coder-7B achieves 42.2% avg on code generation benchmarks.
23.5%
51
arxiv · arxiv/2604.07864 · 2026-04-09
RLVR code training without ground-truth tests. ZeroCoder+DyB4 on Qwen2.5-Coder-7B achieves 42.2% avg on code generation benchmarks.
21.7%
52
arxiv · arxiv/2604.07864 · 2026-04-09
RLVR code training without ground-truth tests. ZeroCoder+DyB4 on Qwen2.5-Coder-7B achieves 42.2% avg on code generation benchmarks.
15.1%
53
arxiv · arxiv/2604.07864 · 2026-04-09
RLVR code training without ground-truth tests. ZeroCoder+DyB4 on Qwen2.5-Coder-7B achieves 42.2% avg on code generation benchmarks.
4.8%