Kimi Code Bench V2

In-house benchmark by Moonshot AI designed to evaluate coding agents on realistic tasks. It has diverse software engineering tasks across 10+ mainstream programming languages and a full production tech stack covering tasks from internal engineering use cases, production incidents, and real-world open-source projects, with emphasis on backend services, infrastructure, performance engineering, systems programming, security, frontend development, and ML/data engineering.

Category

Agentic Coding

Max Score

100.0

Score Type

percent

Active

Yes

Model Scores

Scores for Kimi Code Bench V2

Model Score Date Verified
Kimi K2.6
50.90
23.08.2026 Verified
Kimi K2.7 Code
61.33%
23.08.2026 Verified
GPT-5.5
100.00%
23.08.2026 Verified
Claude Opus 4.8
91.16%
23.08.2026 Verified