Kimi Code Bench V2
In-house benchmark by Moonshot AI designed to evaluate coding agents on realistic tasks. It has diverse software engineering tasks across 10+ mainstream programming languages and a full production tech stack covering tasks from internal engineering use cases, production incidents, and real-world open-source projects, with emphasis on backend services, infrastructure, performance engineering, systems programming, security, frontend development, and ML/data engineering.
Category
Agentic Coding
Max Score
100.0
Score Type
percent
Active
Yes
Model Scores
Scores for Kimi Code Bench V2
| Model | Score | Date | Verified |
|---|---|---|---|
| Kimi K2.6 |
50.90
|
23.08.2026 | Verified |
| Kimi K2.7 Code |
61.33%
|
23.08.2026 | Verified |
| GPT-5.5 |
100.00%
|
23.08.2026 | Verified |
| Claude Opus 4.8 |
91.16%
|
23.08.2026 | Verified |