AI Coding Models Leaderboard
Aggregated official benchmark mirrors and real-world coding task scenarios. 100% verified, objective, and zero paid placement.
π 7 official boards Β· mirrored 1:1
Each card previews a provider's Top 5. Click "View full" for the complete board, or jump to the official page. Data is 100% from each provider β we do no aggregation or math.
Agent ScoreΒ· 33 models- #1Claude Opus 512.47
- #2Claude Fable 511.57
- #3Kimi K310.41
- #4GPT-5.6 Sol9.74
- #5Claude Opus 4.89.55
Pass rate%Β· 23 models- #1Claude Fable 572.9%
- #2Grok 4.669.9%
- #3Gemini 3.7 Flash68.4%
- #4GPT-5.6 Sol67.2%
- #5Grok 4.566.7%
AA IndexΒ· 80 models- #1Claude Opus 563.1
- #2Claude Fable 562.1
- #3GPT-5.6 Sol60.9
- #4Grok 4.660.9
- #5Kimi K359.7
Avg score%Β· 32 models- #1Claude Fable 583.4%
- #2GPT-5.6 Sol81.6%
- #3GPT-5.580.8%
- #4Claude Opus 580.5%
- #5Gemini 3.7 Flash79.9%
Pass rate%Β· 35 models- #1GPT-5.6 Sol88.0%
- #2o3-pro84.9%
- #3Gemini 2.5 Pro83.1%
- #4o381.3%
- #5Grok 4.679.6%
Pass@1%Β· 15 models- #1DeepSeek-V3100.0%
- #2GPT-4o100.0%
- #3GPT-4o mini100.0%
- #4OpenAI o3-mini (High Reasoning Coding)100.0%
- #5o4-mini100.0%
HumanEval+%Β· 21 models- #1o189.0%
- #2o1-mini89.0%
- #3Qwen3.6-Plus87.2%
- #4GPT-4o87.2%
- #5DeepSeek-V386.6%
LMArena AgentAgent
| # | Model | Agent Score | Value Score |
|---|---|---|---|
| #1 | Claude Opus 5Anthropic | 12.47 | 5.6 |
| #2 | Claude Fable 5Anthropic | 11.57 | 7.9 |
| #3 | Kimi K3Moonshot AI | 10.41 | 24.7 |
| #4 | GPT-5.6 SolOpenAI | 9.74 | 11.6 |
| #5 | Claude Opus 4.8Anthropic | 9.55 | 14.7 |
| #6 | GPT-5.5OpenAI | 8.51 | 17.7 |
| #7 | Claude Sonnet 5Anthropic | 6.62 | 35 |
| #8 | DeepSeek-V4-ProDeepSeek | 6.26 | 139.5 |
| #9 | Qwen3.8-MaxQwen | 6.2 | 47.5 |
| #10 | Grok 4.5xAI | 6.17 | 46.9 |
| #11 | GLM-5.2Zhipu AI | 5.82 | 61 |
| #12 | GPT-5.4OpenAI | 4.98 | 32.9 |
| #13 | GPT-5.6 LunaOpenAI | 4.04 | 291.7 |
| #14 | DeepSeek-V4-FlashDeepSeek | 3.99 | 395.2 |
| #15 | Gemini 3.7 FlashGoogle | 3.32 | 92.7 |
| #16 | GPT-5.6 TerraOpenAI | 3.19 | 30.5 |
| #17 | Claude Sonnet 4.6Anthropic | 2.88 | 21.2 |
| #18 | Kimi K2.7 CodeMoonshot AI | 0.43 | 53.4 |
| #19 | GLM-5.1Zhipu AI | 0 | 52.9 |
| #20 | Gemini 3.5 FlashGoogle | -0.68 | 35.5 |
| #21 | Qwen3.7-MaxQwen | -0.75 | 70.4 |
| #22 | Gemini 3.1 ProGoogle | -0.88 | 58.2 |
| #23 | Kimi K2.6Moonshot AI | -1.1 | 68.9 |
| #24 | Gemini 3.6 FlashGoogle | -3.04 | 83.3 |
| #25 | MiniMax-M3MiniMax | -3.15 | 218.4 |
| #26 | Mistral Small 4Mistral AI | -6.82 | 633 |
| #27 | Mistral Medium 3.5Mistral AI | -7.6 | 96.8 |
| #28 | Grok 4.3xAI | -9.02 | 60.2 |
| #29 | -9.48 | 79.2 | |
| #30 | Gemini 2.5 ProGoogle | -10.82 | 37.3 |
| #31 | Gemini 3.5 Flash-LiteGoogle | -10.82 | 106.7 |
| #32 | MiniMax-M2.7MiniMax | -11.86 | 147.8 |
| #33 | Nemotron-UltraNvidia | -15.36 | 3.3 |
The "Cross-board ranks" column shows this model's rank on the other boards at a glance (grey = not tested there). Boards use different methodologies, so raw values are not directly comparable.
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 β we recompute nothing.
π 7 benchmark providers
LMArena Agent board: real agentic coding sessions scored by human blind votes (signal-based agent score). Sourced from arena.ai/leaderboard/agent and mirrored 1:1 β we recompute nothing.
Agent Score33 modelsCursorBench is Cursor's first-party agentic coding benchmark (v3.2). Tasks come from real Cursor sessions with ambiguous, multi-file requirements. Metric: pass rate (%). Mirrored from Cursor's official cursor.com/cursorbench snapshot.
Pass rate%23 modelsArtificial Analysis aggregates 20+ benchmarks (MMLU-Pro, GPQA, HLE, SciCode, IFBench, Terminal-bench, etc.) into a single Intelligence Index β the most comprehensive commercial evaluator for "overall capability".
AA Index80 modelsLiveBench refreshes its questions monthly to prevent contamination. Covers reasoning, math, coding, language, instruction following, and data analysis.
Avg score%32 models133 real coding tasks (bug-fix / feature) across 6 languages. Metric: pass_rate_2.
Pass rate%35 modelsContinuously updated competitive-programming problems. Metric: Pass@1.
Pass@1%15 modelsAll rankings and numbers come directly from official sources; we only do name matching and aggregate presentation, no weighting or re-ranking. Boards use different methodologies, so do not compare raw values across boards.