Kimi K2.7 Code

Moonshot AI

Parameters

1.0T total / 32.0B active

MoE: total / active

Architecture

Mixture-of-Experts (MoE)

Released

11.06.2026

License

Modified MIT License (Kimi)

Open Weights Commercial Use Multimodal BF16/F32/I32 (native INT4 quantization available) Kimi English Chinese

Input Modalities

text image video

Output Modalities

text

Context (native)

262,144 tokens

Context (extended)

262,144 tokens

Openness Index Score 70.0/100

About

Kimi K2.7 Code (moonshotai/Kimi-K2.7-Code) is Moonshot AI's coding-focused agentic model built upon Kimi K2.6 - 1T total parameters with 32B activated per token, released June 11, 2026 under the Modified MIT License. Architecture: Mixture-of-Experts with 61 layers (including 1 dense layer), attention hidden dimension 7168, MoE per-expert dimension 2048, 64 attention heads, 384 experts with 8 selected + 1 shared per token, Multi-head Latent Attention (MLA), SwiGLU activation, 160K vocabulary, a 400M-parameter MoonViT vision encoder, and a 256K-token context.

It delivers substantial improvements on real-world long-horizon coding tasks, strengthening end-to-end task completion across complex software engineering workflows, and improves token efficiency by reducing thinking-token usage approximately 30% compared with Kimi K2.6. The model ships with native INT4 quantization (the same method as Kimi-K2-Thinking), and forces thinking with preserve_thinking mode so reasoning context persists across tool calls.

Training Data Coding-focused agentic model built upon Kimi K2.6. Uses native INT4 quantization (same method as Kimi-K2-Thinking). Forces thinking and preserve_thinking mode.

Benchmark Scores

Benchmark Score Date
Kimi Code Bench V2
coding_agent
61.33%
23.08.2026
Program Bench
coding_agent
25.48%
23.08.2026
MLS-Bench-Lite
coding_agent
36.21%
23.08.2026
Kimi Claw 24/7 Bench
agentic
40.40%
23.08.2026
MCP-Atlas
general_agent
86.73%
23.08.2026
MCPMark-Verified
agentic
41.29%
23.08.2026
WildClawBench
coding_agent
62.35%
23.08.2026
LHTB Solved
coding_agent
100.00%
23.08.2026

API Usage Examples

Kimi K2.7 Code supports OpenAI-compatible API. Key usage patterns:\n\n1. Simple Chat: Standard chat.completions.create() with reasoning content in response.\n2. Chat with Image: Pass image\u005furl with base64-encoded images.\n3. Chat with Video: Pass video\u005furl with base64-encoded videos (experimental, official API only).\n4. Preserve Thinking: Full reasoning content is retained across multi-turn interactions.\n5. Interleaved Thinking + Multi-Step Tool Call: See K2 Thinking documentation.\n\nFor detailed code examples, see the HuggingFace model card.

Key Features

  1. Coding-focused agentic model built upon Kimi K2.6\n2. 30% reduction in thinking-token usage compared to K2.6\n3. Native INT4 quantization (same method as Kimi-K2-Thinking)\n4. Forced thinking + preserve\u005fthinking mode for enhanced multi-turn coding agent performance\n5. Multimodal: Supports text, image, and video input\n6. MLA attention mechanism with SwiGLU activation\n7. MoonViT vision encoder (400M params)\n8. Kimi Code CLI as recommended coding agent framework (https://www.kimi.com/code)\n9. Open source under Modified MIT License\n10. Available via OpenAI/Anthropic-compatible API on platform.moonshot.ai

Context Length and Modalities

Context Length: 256K (262,144 tokens)\nInput Modalities: Text, Image, Video\nOutput Modalities: Text\n\nNote: Chat with video content is an experimental feature and is only supported in the official API for now (not in third-party APIs deployed with vLLM or SGLang).

Thinking and Preserve Thinking Mode

Kimi K2.7 Code forces thinking and preserve\u005fthinking as True. This feature is enabled by default and cannot be disabled.\n\nPreserve thinking retains full reasoning content across multi-turn interactions and enhances performance in coding agent scenarios.\n\nRecommended settings for Thinking mode:\n- Temperature: 1.0\n- Top-p: 0.95\n- Instant mode is not supported\n\nK2.7-Code also supports Interleaved Thinking and Multi-Step Tool Call (same design as K2 Thinking).

Deployment

Kimi-K2.7-Code API is available on https://platform.moonshot.ai with OpenAI/Anthropic-compatible API.\n\nRecommended inference engines:\n- vLLM\n- SGLang\n- KTransformers\n\nKimi-K2.7-Code has the same architecture as Kimi-K2.5/Kimi-K2.6, and the deployment method can be directly reused. The version requirement for transformers is >=4.57.1, <5.0.0.\n\nDeployment examples: https://huggingface.co/moonshotai/Kimi-K2.7-Code/blob/main/docs/deploy_guidance.md

Model Architecture

Architecture: Mixture-of-Experts (MoE) Total Parameters: 1T Activated Parameters: 32B Number of Layers (Dense layer included): 61 Number of Dense Layers: 1 Attention Hidden Dimension: 7168 MoE Hidden Dimension (per Expert): 2048 Number of Attention Heads: 64 Number of Experts: 384 Selected Experts per Token: 8 Number of Shared Experts: 1 Vocabulary Size: 160K Context Length: 256K Attention Mechanism: MLA Activation Function: SwiGLU Vision Encoder: MoonViT (400M parameters)

Kimi K2.7 Code Overview

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

Architecture

Decoder Block input Embedding vocab 164K · d 7168 Full Attention MLA · 64 heads ×61 MoE FFN 384 experts · top-8 · +1 shared · dᴻ 2048 Final Norm LM Head vocab 164K output
Attention
Multi-head Latent Attention
MoE
384 experts · top-8 per token
Layers
61
Hidden size
7168
Context
262K tokens
RoPE θ
50K
Parameters
1000000M
Active params
32000M

Source: Hugging Face config.json · KimiK25ForConditionalGeneration · model repo

Type: Mixture-of-Experts (MoE)
Attention: MLA
Decoder: Transformer
MoE: yes (384 experts)
Routing: top-8 with 1 shared expert
Layers 61
Context length 262K
Experts 384
Shared experts 1
Attention heads 64
Attention hidden size 7168
Vocabulary 160K
Activation SwiGLU
Vision encoder MoonViT
Moe Hidden Dim Per Expert 2048
Num Dense Layers 1
Quantization native INT4 (same as Kimi-K2-Thinking)
Selected Experts Per Token 8
Vision encoder params 400M

Training Pipeline

  1. 1
    other

    Coding-Focused Fine-Tuning from Kimi K2.6

    Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6. Uses native INT4 quantization (same method as Kimi-K2-Thinking). Forces thinking and preserve_thinking mode.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Coding-focused post-training corpus (built on Kimi K2.6) rl — —

Linked Resources

Trend Analysis

24h Change

+0.2%

7d Change

+2.0%

Current

18,018

huggingface

downloads

+0.4%

huggingface

likes

+0.0%

ollama

downloads

+0.3%

huggingface

downloads_all_time

+0.2%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 229,700 pulls daily 01.09.2026
huggingface downloads_all_time 1,788,568 daily 01.09.2026
huggingface followers 18,018 daily 01.09.2026
huggingface likes 1,376 daily 01.09.2026
huggingface downloads 230,265 daily 01.09.2026
ollama downloads 229,000 pulls daily 31.08.2026
huggingface downloads_all_time 1,785,087 daily 31.08.2026
huggingface followers 17,977 daily 31.08.2026
huggingface likes 1,376 daily 31.08.2026
huggingface downloads 229,309 daily 31.08.2026
ollama downloads 228,200 pulls daily 30.08.2026
huggingface downloads_all_time 1,783,855 daily 30.08.2026
huggingface followers 17,928 daily 30.08.2026
huggingface likes 1,375 daily 30.08.2026
huggingface downloads 250,481 daily 30.08.2026
ollama downloads 227,500 pulls daily 29.08.2026
huggingface followers 17,884 daily 29.08.2026
huggingface likes 1,375 daily 29.08.2026
huggingface downloads 276,434 daily 29.08.2026
ollama downloads 226,800 pulls daily 28.08.2026
huggingface followers 17,836 daily 28.08.2026
huggingface likes 1,375 daily 28.08.2026
huggingface downloads 299,848 daily 28.08.2026
ollama downloads 226,000 pulls daily 27.08.2026
huggingface followers 17,787 daily 27.08.2026
huggingface likes 1,372 daily 27.08.2026
huggingface downloads 334,742 daily 27.08.2026
ollama downloads 225,200 pulls daily 26.08.2026
huggingface followers 17,709 daily 26.08.2026
huggingface likes 1,372 daily 26.08.2026
huggingface downloads 363,955 daily 26.08.2026
ollama downloads 224,400 pulls daily 25.08.2026
huggingface followers 17,658 daily 25.08.2026
huggingface likes 1,371 daily 25.08.2026
huggingface downloads 390,528 daily 25.08.2026
ollama downloads 223,700 pulls daily 24.08.2026
huggingface followers 17,609 daily 24.08.2026
huggingface likes 1,371 daily 24.08.2026
huggingface downloads 412,917 daily 24.08.2026
huggingface followers 17,556 daily 23.08.2026
huggingface likes 1,370 daily 23.08.2026
huggingface downloads 440,001 daily 23.08.2026

View full metric history →

Related Models