Parameters
2.8T total / 104.0B active
MoE: total / active
Architecture
Mixture-of-Experts (MoE)
Released
13.06.2026
License
Kimi K3 License
Input Modalities
Output Modalities
Context (native)
1,048,576 tokens
Context (extended)
1,048,576 tokens
About
Kimi K3 (moonshotai/Kimi-K3, released June 13, 2026 under the Kimi K3 License) is Moonshot AI's open-weight, native multimodal agentic model and its most capable model to date - a 2.8T-parameter MoE model (104B activated per token, 93 layers: 1 dense + 69 KDA + 24 Gated MLA attention layers; activates 16 of 896 experts), the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work and reasoning.
It is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and scales up MoE sparsity with a Stable LatentMoE framework yielding an approximate 2.5x improvement in overall scaling efficiency over Kimi K2. Kimi K3 understands text, images and video in the same model, with a 1M-token context window, a MoonViT-V2 vision encoder (401M parameters), SiTU-GLU activation, and native MXFP4 quantization (quantization-aware training from the SFT stage onward: MXFP4 weights / MXFP8 activations for broad hardware compatibility). It sustains long engineering sessions with minimal oversight - navigating massive repositories and orchestrating terminal tools from GPU kernels and compilers to vision-in-the-loop game dev, CAD and chip design - and advances end-to-end agentic knowledge work with interactive visualizations and deep research. Full weights are open for research, deployment and further innovation.
Training Data Quantization-aware training from SFT stage onward using MXFP4 weights with MXFP8 activations
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
GPQA Diamond
stem_reasoning
|
98.49%
|
18.08.2026 |
|
CritPt (no tools)
stem_reasoning
|
71.92%
|
18.08.2026 |
|
AA-LCR
long_context
|
93.38%
|
18.08.2026 |
|
HLE (with tools)
stem_reasoning
|
81.99%
|
18.08.2026 |
|
DeepSWE 1.1
coding_agent
|
89.88%
|
18.08.2026 |
|
FrontierSWE
coding_agent
|
88.18%
|
18.08.2026 |
|
MLS-Bench-Lite
coding_agent
|
93.10%
|
18.08.2026 |
|
PostTrainBench
coding_agent
|
80.20%
|
18.08.2026 |
|
Terminal Bench 2.1
coding_agent
|
97.45%
|
18.08.2026 |
|
SciCode (subtask)
stem_reasoning
|
99.03%
|
18.08.2026 |
|
WildClawBench
coding_agent
|
73.53%
|
18.08.2026 |
|
BrowseComp
general_agent
|
100.00%
|
18.08.2026 |
|
DeepSearch QA
general_agent
|
100.00%
|
18.08.2026 |
|
MCP-Atlas
general_agent
|
97.95%
|
18.08.2026 |
|
MCPMark
general_agent
|
100.00%
|
18.08.2026 |
|
Automation-Bench
general_agent
|
45.45%
|
18.08.2026 |
|
JobBench
general_agent
|
70.35%
|
18.08.2026 |
|
Agents' Last Exam
general_agent
|
83.49%
|
18.08.2026 |
|
Apex-Agents
general_agent
|
97.79%
|
18.08.2026 |
|
OfficeQA Pro
general_agent
|
92.62%
|
18.08.2026 |
|
SpreadSheetBench-v1
general_agent
|
34.80
|
18.08.2026 |
|
Toolathlon Verified
general_agent
|
97.21%
|
18.08.2026 |
|
TauBench V3 Banking
general_agent
|
100.00%
|
18.08.2026 |
|
GDPVal-AA v2
general_agent
|
92.08%
|
18.08.2026 |
|
MMMU-Pro
vision_language
|
100.00%
|
18.08.2026 |
|
CharXiv (RQ)
document_understanding
|
100.00%
|
18.08.2026 |
|
MathVision
vision_language
|
100.00%
|
18.08.2026 |
|
ZEROBench
vision_language
|
82.61%
|
18.08.2026 |
|
MMVU
video_understanding
|
100.00%
|
18.08.2026 |
|
BabyVision
vision_language
|
94.94%
|
18.08.2026 |
|
VideoMME (w sub.)
video_understanding
|
100.00%
|
18.08.2026 |
|
Humanity's Last Exam
stem_reasoning
|
78.41%
|
13.08.2026 |
|
Cybergym
general_agent
|
83.60%
|
13.08.2026 |
|
DeepSWE
coding_agent
|
92.85%
|
13.08.2026 |
|
DSBench-FullStack
coding_agent
|
91.29%
|
13.08.2026 |
|
DSBench-Hard
coding_agent
|
81.05%
|
13.08.2026 |
|
Terminal-Bench 3.0
coding_agent
|
42.67%
|
28.08.2026 |
|
NL2Repo
coding_agent
|
73.38%
|
28.08.2026 |
|
ProgramBench
coding_agent
|
21.04%
|
28.08.2026 |
|
SWE-Marathon
coding_agent
|
96.12%
|
28.08.2026 |
|
ExploitGym (2h)
cybersecurity
|
10.89%
|
28.08.2026 |
|
ExploitGym (6h)
cybersecurity
|
16.48%
|
28.08.2026 |
|
ExploitBench
cybersecurity
|
14.55%
|
28.08.2026 |
|
SWE-bench Multilingual
coding_agent
|
89.70%
|
28.08.2026 |
|
SWE-bench Pro
coding_agent
|
79.12%
|
28.08.2026 |
|
SWE Atlas - QnA
coding_agent
|
47.45%
|
28.08.2026 |
|
SWE Atlas - TW
coding_agent
|
48.52%
|
28.08.2026 |
|
SWE Atlas - RF
coding_agent
|
59.43%
|
28.08.2026 |
|
Terminal-Bench 2.1 (Terminus-2)
coding_agent
|
99.26%
|
28.08.2026 |
|
Hy-Backend 2.0 (Internal)
coding_agent
|
50.00%
|
28.08.2026 |
|
Hy-SWE Max Verified (Internal)
coding_agent
|
78.67%
|
28.08.2026 |
|
Hy-CompanyBench V2 (Internal)
general_agent
|
78.09%
|
28.08.2026 |
|
WideSearch
general_agent
|
89.63%
|
28.08.2026 |
|
$OneMillion-Bench
general_capabilities
|
82.51%
|
28.08.2026 |
|
DRACO
general_agent
|
52.56%
|
28.08.2026 |
|
Hy-LifeSearch (Internal)
general_agent
|
28.16%
|
28.08.2026 |
|
Hy-BrowseComp-Pro2 (Internal)
general_agent
|
78.38%
|
28.08.2026 |
|
WorkSpaceBench
general_agent
|
40.48%
|
28.08.2026 |
|
BankerToolBench
general_agent
|
53.33%
|
28.08.2026 |
|
E-Bench (Internal)
general_agent
|
78.82%
|
28.08.2026 |
|
E-Bench-Code (Internal)
coding_agent
|
70.00%
|
28.08.2026 |
|
Hy-FinAgentBench (Internal)
domain_finance
|
66.67%
|
28.08.2026 |
|
Hy-FinmodelBench v2 (Internal)
domain_finance
|
63.64%
|
28.08.2026 |
|
BioMysteryBench
stem_reasoning
|
35.16%
|
28.08.2026 |
|
SUPERChem
stem_reasoning
|
59.34%
|
28.08.2026 |
|
ArXivMath
stem_reasoning
|
32.73%
|
28.08.2026 |
|
HorizonMath (pass@4)
stem_reasoning
|
50.28%
|
28.08.2026 |
|
MathArena Apex 2025
stem_reasoning
|
56.36%
|
28.08.2026 |
|
BrokenArXiv
stem_reasoning
|
58.04%
|
28.08.2026 |
|
SkillsBench Avg5
coding_agent
|
76.24%
|
28.08.2026 |
|
OSWorld-Verified
general_agent
|
100.00%
|
18.08.2026 |
Kimi K3 Evaluation Methodology
Evaluation Methodology
General Settings
- All Kimi K3 results:
reasoning_effort="max",temperature=1.0 - Single-step tasks (GPQA Diamond, HLE-Full, vision benchmarks without tools):
top-p=0.95 - Agentic tasks:
top-p=1.0 - For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision, ZeroBench: each cell reports scores without and with tool augmentation (general tools for HLE-Full, Python for vision benchmarks)
Reasoning & Knowledge Benchmarks
- CritPt and AA-LCR: Scores cited from Artificial Analysis as of July 23, 2026
Coding Benchmarks
- DeepSWE: Kimi K3 evaluated with Kimi Code harness (67.3 with mini-SWE-agent harness); GLM-5.2 score from GLM-5.2 release blog; others from official DeepSWE leaderboard
- Terminal-Bench 2.1: Kimi K3 with Kimi Code harness; other models report best score across harnesses
- ProgramBench: Kimi K3 with Kimi Code harness; GLM-5.2 from release blog; others from Vals AI
- SWE-Marathon: Kimi K3, Claude Opus 4.8, Claude Fable 5 with Claude Code harness; GPT-5.6 Sol with Codex harness; evaluation on H20-calibrated branch
- FrontierSWE: Kimi K3 with Kimi Code harness; GPT-5.6 Sol with Codex harness; others from FrontierSWE
- PostTrainBench: Kimi K3, Claude Fable 5 with Claude Code harness, GPT-5.6 Sol with Codex harness; averaged over 3 runs on H20 GPUs
- MLS-Bench-Lite: Kimi K3 with Kimi Code harness; Claude models with Claude Code; GPT models with Codex
- SciCode: Scores from Artificial Analysis (July 23, 2026)
- Kimi Code Bench 2.0: In-house benchmark; Kimi K3 with Kimi Code harness (73.7 with Claude Code harness)
Agentic Benchmarks
- BrowseComp: Context-compaction at 300K tokens; full 1M context scores 90.4
- OfficeQA Pro: PDFs rendered as images, no machine-readable text
- MCP-Atlas: 500-task public subset, 100-turn limit, Gemini 3.1 Pro as judge
- AutomationBench: 600-task public subset, official GitHub setup
- GDPval-AA v2, AA-Briefcase, τ³-Banking, Harvey Lab-AA, APEX-Agents: From Artificial Analysis and APEX-Agents leaderboard (July 23, 2026)
- CorpFin v2, Finance Agent v2, Legal Research Bench: From Vals AI
- Agents' Last Exam: From official leaderboard (July 23, 2026); each model paired with specific harness
Multimodal Benchmarks
- Except ZeroBench (5 runs, official setting), all multimodal scores averaged over 3 runs
- MMMU-Pro: official protocol, original input order, images prepended to text input
- PerceptionBench: In-house benchmark for atomic visual perception
Kimi K3 Citation and Resources
Citation and Resources
Primary Sources
| Resource | URL |
|---|---|
| Tech Blog | kimi.com/blog/kimi-k3 |
| Full Technical Report (PDF) | github.com/MoonshotAI/Kimi-K3/k3_tech_report.pdf |
| HuggingFace Model Page | huggingface.co/moonshotai/Kimi-K3 |
| API Platform | platform.kimi.ai |
Organization Links
| Resource | URL |
|---|---|
| Moonshot AI Website | moonshot.ai |
| Kimi Product | kimi.com |
| Moonshot AI on HuggingFace | huggingface.co/moonshotai |
| twitter.com/kimi_moonshot | |
| Discord | discord.gg/TYU2fdJykW |
| ModelScope | modelscope.cn/organization/moonshotai |
Contact
- Email: support@moonshot.ai
Quickstart Guides
| Guide | URL |
|---|---|
| Kimi K3 Quickstart | platform.kimi.ai/docs/guide/kimi-k3-quickstart |
| Thinking Effort Guide | platform.kimi.ai/docs/guide/use-thinking-effort |
| Kimi Code CLI | kimi.com/code |
Kimi K3 Thinking Control and Reasoning Effort
Thinking Control
Always-On Thinking
Kimi K3 always has thinking enabled and returns reasoning_content in every response.
Reasoning Effort Levels
The reasoning_effort request field controls thinking depth:
| Level | Description |
|---|---|
"low" |
Minimal reasoning, faster responses |
"high" |
Moderate reasoning depth |
"max" (default) |
Maximum reasoning depth |
Preserved Thinking History
Kimi K3 was trained in preserved thinking history mode. This means:
- For multi-turn conversations, the complete assistant message (including
reasoning_contentandtool_calls) must be passed back tomessagesas-is. - Only passing
content(withoutreasoning_content) will degrade performance in multi-turn settings. - This is critical for tool calls — the thinking context must be preserved across turns.
Evaluation Settings
- All benchmark results use
reasoning_effort="max"withtemperature=1.0 - Single-step tasks:
top-p=0.95 - Agentic tasks:
top-p=1.0
Additional Resources
Kimi K3 Training and Quantization Approach
Training Approach
Quantization-Aware Training (QAT)
Kimi K3 applies quantization-aware training from the SFT stage onward, using:
- MXFP4 weights (4-bit microscaling float)
- MXFP8 activations (8-bit microscaling float)
This means the model was trained with the quantization format in the loop, rather than being quantized post-hoc. This preserves model quality at reduced precision.
Architecture Training Stability
The Stable LatentMoE framework enables stable training of the high-sparsity MoE architecture (896 experts, 16 activated). This framework was key to achieving the 2.5× scaling efficiency improvement over Kimi K2.
Training Data
- Languages: English, Chinese
- Training data info: Quantization-aware training from SFT stage onward using MXFP4 weights with MXFP8 activations
- Full details available in the Technical Report (PDF)
Kimi K3 Vision and Multimodal Capabilities
Vision and Multimodal Capabilities
Native Multimodality
Kimi K3 understands text, images, and video within the same model, powered by its native multimodal architecture.
Vision Encoder
- Encoder: MoonViT-V2
- Parameters: 401M
- Supports text, image, and video inputs
Modalities
| Direction | Modalities |
|---|---|
| Input | Text, Image, Video |
| Output | Text |
Vision Benchmark Performance
| Benchmark | Kimi K3 (max) |
|---|---|
| WorldVQA ForceAnswer | 51.0 |
| OmniDocBench | 91.1 |
| PerceptionBench | 58.5 |
| Video-MME (w. sub) | 90.0 |
| MMVU | 82.1 |
| BabyVision w/ python | 85.7 |
| MMMU-Pro | 81.6 / 83.4 |
| CharXiv (RQ) | 84.8 / 91.3 |
| MathVision | 94.3 / 97.8 |
| ZeroBench (pass@5) | 23.0 / 41.0 |
Notes
- Except for ZeroBench (5 runs, official setting), all multimodal scores are averaged over 3 runs.
- MMMU-Pro is evaluated following the official protocol, preserving the original input order and prepending images to the text input.
- PerceptionBench is an in-house benchmark focusing on atomic visual perception capabilities.
Kimi K3 Model Ecosystem
Model Ecosystem
Model Tree
| Type | Count | Link |
|---|---|---|
| Adapters | 1 model | HF Models |
| Finetunes | 43 models | HF Models |
| Quantizations | 47 models | HF Models |
Quantization Compatibility
The model page lists compatibility with:
- llama.cpp
- LM Studio
- Jan
- Ollama
Community
- 171 discussions on HuggingFace
- 46 Spaces using moonshotai/Kimi-K3
- 2,787,971 downloads last month
- 11 likes
- 17,600 followers for Moonshot AI
Collection
Part of the Kimi K3 Collection — 2 items, 113 likes.
Kimi K3 Key Highlights
Key Highlights
Kimi K3 represents several breakthroughs as Moonshot AI's most capable model:
- World's first open 3T-class model — 2.8T parameters with open weights
- Novel KDA Architecture — Kimi Delta Attention + Attention Residuals, a departure from standard attention mechanisms
- 2.5× scaling efficiency over Kimi K2 via Stable LatentMoE framework (16/896 experts activated)
- Native multimodality — text, image, and video understanding in a single model
- 1M-token context window — native, no extension needed
- Long-horizon coding — operates with minimal human oversight across massive repositories
- Agentic knowledge work — deep research with interactive visualizations, dashboards, motion design
- Native MXFP4 quantization — quantization-aware training from SFT stage for broad hardware compatibility
- Open frontier weights — full weights under Kimi K3 License for research and commercial deployment
Kimi K3 Context Length
Context Length
| Property | Value |
|---|---|
| Native Context Length | 1,048,576 tokens (1M) |
| Extended Context Length | 1,048,576 tokens (1M) |
Kimi K3 supports a 1-million-token context window natively. The model can process text, images, and video within this context.
Context Management
- For BrowseComp evaluation, a context-compaction strategy is triggered at 300K tokens.
- When evaluated with the full 1M-token context window and no context management, Kimi K3 achieves a BrowseComp score of 90.4 (vs. 91.2 with compaction).
Kimi K3 License Information
License
Both the code repository and the model weights are released under the Kimi K3 License.
- The model is open-weight with commercial use allowed.
- The license terms are available at the above URL.
Kimi K3 Coding Agent Framework
Coding Agent Framework
Kimi K3 works best with Kimi Code CLI as its agent framework. Users can run Kimi Code in their terminal and select Kimi K3 using the /model command.
Capabilities
- Long-Horizon Coding: Sustains long engineering sessions with minimal human oversight
- Repository Navigation: Navigates massive codebases effectively
- Terminal Tool Orchestration: From GPU kernel optimization and compiler development to vision-in-the-loop game dev, CAD, and chip design
- Multi-tool Coordination: Coordinates terminal tools, file editors, and build systems across complex workflows
Kimi K3 API Usage and Thinking Mode
Model Usage
Thinking Mode
Kimi K3 always has thinking enabled and will return reasoning_content. Thinking effort is configured with the top-level reasoning_effort request field:
"low"— minimal reasoning"high"— moderate reasoning"max"(default) — maximum reasoning
Preserved Thinking History Mode
Kimi K3 was trained in the preserved thinking history mode. For multi-turn conversations and tool calls, the complete assistant message returned by the API must be passed back to messages as-is — including reasoning_content and tool_calls, not just content.
Example: Multi-turn with Preserved Thinking
import openai
def chat_with_preserved_thinking(client: openai.OpenAI, model_name: str):
messages = [
{
"role": "user",
"content": "Tell me three random numbers."
},
{
"role": "assistant",
"reasoning_content": "I'll start by listing five numbers: 473, 921, 235, 215, 222, and I'll tell you the first three.",
"content": "473, 921, 235"
},
{
"role": "user",
"content": "What are the other two numbers you have in mind?"
}
]
response = client.chat.completions.create(
model=model_name,
messages=messages,
stream=False,
max_tokens=4096,
reasoning_effort="max",
)
# the assistant should mention 215 and 222 from prior reasoning content
print(f"response: {response.choices[0].message.reasoning}")
return response.choices[0].message.content
Additional Resources
- Kimi K3 Quickstart Guide
- Thinking Effort Guide
- Full guides cover: vision input, structured output, partial mode, tool choice, dynamic tool loading, context caching.
Kimi K3 Deployment Options
Deployment
API Access
Kimi K3's API is available on https://platform.kimi.ai by selecting kimi-k3. The API provides OpenAI/Anthropic-compatible endpoints.
Supported Inference Engines
| Engine | Link | Recipes |
|---|---|---|
| vLLM | github.com/vllm-project/vllm | recipes.vllm.ai/moonshotai/Kimi-K3 |
| SGLang | github.com/sgl-project/sglang | docs.sglang.io/cookbook/.../Kimi-K3 |
| TokenSpeed | github.com/lightseekorg/tokenspeed | lightseek.org/tokenspeed/recipes |
Kimi K3 Native MXFP4 Quantization
Native MXFP4 Quantization
Kimi K3 applies quantization-aware training (QAT) from the SFT stage onward, using:
- MXFP4 weights — 4-bit microscaling float format for model weights
- MXFP8 activations — 8-bit microscaling float format for activations
This approach ensures broad hardware compatibility while maintaining model quality. The quantization is native (baked into the model during training), not a post-hoc compression step, which preserves the model's reasoning and coding capabilities at the reduced precision.
Tensor Types in Released Weights
- F32
- BF16
- U8
The model is distributed in Safetensors format with the native MXFP4 quantization already applied.
Kimi K3 Benchmark Results
Evaluation Results
All Kimi K3 results obtained with reasoning effort set to 'max' and temperature = 1.0. For single-step tasks (GPQA Diamond, HLE-Full, vision benchmarks without tools), top-p = 0.95; for agentic tasks, top-p = 1.0.
Reasoning & Knowledge
| Benchmark | Kimi K3 (max) |
|---|---|
| GPQA Diamond | 93.5 |
| CritPt | 23.4 |
| AA-LCR | 74.7 |
| HLE-Full | 43.5 / 56.0 |
Coding
| Benchmark | Kimi K3 (max) |
|---|---|
| DeepSWE | 67.5 |
| ProgramBench | 77.8 |
| Terminal-Bench 2.1 | 88.3 |
| FrontierSWE | 81.2 |
| SWE-Marathon | 42.0 |
| PostTrainBench | 36.6 |
| MLS-Bench-Lite | 48.3 |
| SciCode | 58.7 |
| Kimi Code Bench 2.0 | 72.9 |
Agentic
| Benchmark | Kimi K3 (max) |
|---|---|
| BrowseComp | 91.2 |
| DeepSearchQA (F1) | 95.0 |
| ResearchRubrics | 76.2 |
| GDPval-AA v2 (Elo) | 1686 |
| Toolathlon-Verified | 76.5 |
| MCPMark-Verified | 94.5 |
| MCP-Atlas | 84.2 |
| AutomationBench | 30.8 |
| JobBench | 54.3 |
| AA-Briefcase (Elo) | 1548 |
| Agents' Last Exam | 28.3 |
| APEX-Agents | 41.0 |
| OfficeQA Pro | 63.3 |
| SpreadsheetBench 2 | 34.8 |
| OSWorld-Verified | 84.8 |
| OSWorld 2.0 | 58.3 |
| SaaS-Bench | 60.1 |
| τ³-Banking | 33.4 |
| Harvey Lab-AA | 94.6 |
| CorpFin v2 | 71.6 |
| Finance Agent v2 | 54.4 |
| Legal Research Bench | 44.2 |
Vision
| Benchmark | Kimi K3 (max) |
|---|---|
| WorldVQA ForceAnswer | 51.0 |
| OmniDocBench | 91.1 |
| PerceptionBench | 58.5 |
| Video-MME (w. sub) | 90.0 |
| MMVU | 82.1 |
| BabyVision w/ python | 85.7 |
| MMMU-Pro | 81.6 / 83.4 |
| CharXiv (RQ) | 84.8 / 91.3 |
| MathVision | 94.3 / 97.8 |
| ZeroBench (pass@5) | 23.0 / 41.0 |
Footnotes
- CritPt and AA-LCR scores cited from Artificial Analysis (July 23, 2026).
- For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision, and ZeroBench, each cell reports scores without and with tool augmentation.
- DeepSWE: Kimi K3 evaluated with Kimi Code harness (67.3 with mini-SWE-agent harness).
- BrowseComp: Context-compaction strategy triggered at 300K tokens; full 1M-token context scores 90.4.
Kimi K3 Architecture Details
Architecture Summary
| Property | Value |
|---|---|
| Architecture | Mixture-of-Experts (MoE) |
| Total Parameters | 2.8T |
| Activated Parameters | 104B |
| Number of Layers | 93 |
| Number of Dense Layers | 1 |
| Attention-Layer Composition | 69 KDA + 24 Gated MLA |
| Attention Hidden Dimension | 7168 |
| Number of Attention Heads | 96 |
| Latent MoE Dimension | 3584 |
| MoE Hidden Dimension (per Expert) | 3072 |
| Number of Experts | 896 |
| Selected Experts per Token | 16 |
| Number of Shared Experts | 2 |
| Vocabulary Size | 160K |
| Context Length | 1,048,576 |
| Attention Mechanism | KDA & Gated MLA |
| Activation Function | SiTU-GLU |
| Vision Encoder | MoonViT-V2 |
| Parameters of Vision Encoder | 401M |
| Quantization | MXFP4 weights / MXFP8 activations (quantization-aware training) |
| Modality | Text, Image |
Key Architectural Innovations
- Kimi Delta Attention (KDA): Novel attention mechanism used in 69 of 93 layers.
- Gated MLA: Used in 24 layers, complementing KDA for a hybrid attention design.
- Attention Residuals (AttnRes): Residual connections across attention layers for improved gradient flow.
- Stable LatentMoE: Framework for stable training of high-sparsity MoE models.
- SiTU-GLU: Custom activation function combining SiLU with a Gated Linear Unit.
Kimi K3 Model Introduction
Model Introduction
Kimi K3 is an open-weight, native multimodal agentic model and Moonshot AI's most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.
Key Features
- New Architecture: Built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), scaling up MoE sparsity with a Stable LatentMoE framework that activates 16 out of 896 experts — yielding an approximate 2.5× improvement in overall scaling efficiency over Kimi K2.
- Long-Horizon Coding: Operates with minimal human oversight, sustains long engineering sessions, navigates massive repositories, and orchestrates terminal tools — from GPU kernel optimization and compiler development to vision-in-the-loop game dev, CAD, and chip design.
- Agentic Knowledge Work: Produces deep research with interactive visualizations, widgets and dashboards, and motion design and video editing, powered by its native multimodal architecture.
- Native Multimodality & Long Context: Understands text, images, and video within the same model; supports a 1-million-token context window.
- Open Frontier Weights: Full model weights released under the Kimi K3 License for research, deployment, and further innovation.
Architecture
- Attention
- Multi-head Latent Attention
- MoE
- 896 experts
- Layers
- 93
- Hidden size
- 7168
- Context
- 1M tokens
- Parameters
- 2800000M
- Active params
- 104000M
Source: Hugging Face config.json · KimiK3ForConditionalGeneration · model repo
Training Pipeline
-
1
pretraining
Pretraining
2.8T parameter MoE model with Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). Uses Stable LatentMoE framework activating 16 out of 896 experts. Native MXFP4 quantization-aware training from SFT stage onward.
-
2
sft
Supervised Fine-Tuning (SFT)
Quantization-aware training applied from SFT stage onward using MXFP4 weights with MXFP8 activations for broad hardware compatibility.
-
3
rl
Reinforcement Learning (RL)
Trained in preserved thinking history mode. Thinking effort supports low, high, and max settings.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| DeepSWE | evaluation | — | — | |
| GPQA Diamond | evaluation | — | — | |
| WildClawBench | evaluation | — | — | |
| MDPBench | evaluation | — | — |
Linked Resources
Kimi K3 Tech Report
https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf
Kimi K3 Tech Blog
https://www.kimi.com/blog/kimi-k3
Kimi-K3 GitHub Repository
https://github.com/MoonshotAI/Kimi-K3
Kimi K3 Quickstart Guide
https://platform.kimi.ai/docs/guide/kimi-k3-quickstart
Thinking Effort Guide
https://platform.kimi.ai/docs/guide/use-thinking-effort
Kimi Platform
https://platform.kimi.ai
Kimi.com
https://www.kimi.com
Moonshot AI
https://www.moonshot.ai
Kimi Code CLI
https://www.kimi.com/code
vLLM Recipes for Kimi-K3
https://recipes.vllm.ai/moonshotai/Kimi-K3
SGLang Cookbook for Kimi-K3
https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3
TokenSpeed GitHub
https://github.com/lightseekorg/tokenspeed
Kimi K3 License
https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE
Kimi K3 Collection
https://huggingface.co/collections/moonshotai/kimi-k3
ModelScope - Moonshot AI
https://modelscope.cn/organization/moonshotai
Trend Analysis
24h Change
+0.2%
7d Change
+2.0%
Current
18,018
downloads_all_time
+3.5%
likes
+0.1%
downloads
+2.1%
downloads
-0.3%
Usage & Social Metrics
| Source | Metric | Value | Period | Recorded |
|---|---|---|---|---|
| ollama | downloads | 68,700 pulls | daily | 01.09.2026 |
| huggingface | downloads_all_time | 3,735,508 | daily | 01.09.2026 |
| huggingface | followers | 18,018 | daily | 01.09.2026 |
| huggingface | likes | 11,128 | daily | 01.09.2026 |
| huggingface | downloads | 2,783,061 | daily | 01.09.2026 |
| ollama | downloads | 67,300 pulls | daily | 31.08.2026 |
| huggingface | downloads_all_time | 3,610,900 | daily | 31.08.2026 |
| huggingface | followers | 17,977 | daily | 31.08.2026 |
| huggingface | likes | 11,116 | daily | 31.08.2026 |
| huggingface | downloads | 2,792,274 | daily | 31.08.2026 |
| ollama | downloads | 65,700 pulls | daily | 30.08.2026 |
| huggingface | downloads_all_time | 3,446,843 | daily | 30.08.2026 |
| huggingface | followers | 17,928 | daily | 30.08.2026 |
| huggingface | likes | 11,099 | daily | 30.08.2026 |
| huggingface | downloads | 2,794,721 | daily | 30.08.2026 |
| ollama | downloads | 64,200 pulls | daily | 29.08.2026 |
| huggingface | followers | 17,884 | daily | 29.08.2026 |
| huggingface | likes | 11,080 | daily | 29.08.2026 |
| huggingface | downloads | 2,701,014 | daily | 29.08.2026 |
| ollama | downloads | 62,000 pulls | daily | 28.08.2026 |
| huggingface | followers | 17,836 | daily | 28.08.2026 |
| huggingface | likes | 11,061 | daily | 28.08.2026 |
| huggingface | downloads | 2,675,145 | daily | 28.08.2026 |
| ollama | downloads | 60,300 pulls | daily | 27.08.2026 |
| huggingface | followers | 17,787 | daily | 27.08.2026 |
| huggingface | likes | 11,036 | daily | 27.08.2026 |
| huggingface | downloads | 2,829,554 | daily | 27.08.2026 |
| ollama | downloads | 58,200 pulls | daily | 26.08.2026 |
| huggingface | followers | 17,709 | daily | 26.08.2026 |
| huggingface | likes | 11,018 | daily | 26.08.2026 |
| huggingface | downloads | 2,921,257 | daily | 26.08.2026 |
| ollama | downloads | 55,600 pulls | daily | 25.08.2026 |
| huggingface | followers | 17,658 | daily | 25.08.2026 |
| huggingface | likes | 10,995 | daily | 25.08.2026 |
| huggingface | downloads | 2,865,293 | daily | 25.08.2026 |
| ollama | downloads | 53,000 pulls | daily | 24.08.2026 |
| huggingface | followers | 17,609 | daily | 24.08.2026 |
| huggingface | likes | 10,972 | daily | 24.08.2026 |
| huggingface | downloads | 2,787,971 | daily | 24.08.2026 |
| huggingface | followers | 17,556 | daily | 23.08.2026 |
| huggingface | likes | 10,949 | daily | 23.08.2026 |
| huggingface | downloads | 2,727,920 | daily | 23.08.2026 |
| huggingface | followers | 17,515 | daily | 22.08.2026 |
| huggingface | likes | 10,924 | daily | 22.08.2026 |
| huggingface | downloads | 2,612,739 | daily | 22.08.2026 |
| huggingface | followers | 17,482 | daily | 21.08.2026 |
| huggingface | likes | 10,913 | daily | 21.08.2026 |
| huggingface | downloads | 2,448,810 | daily | 21.08.2026 |
| huggingface | followers | 17,418 | daily | 20.08.2026 |
| huggingface | likes | 10,883 | daily | 20.08.2026 |