Kimi K3

Moonshot AI

Parameters

2.8T total / 104.0B active

MoE: total / active

Architecture

Mixture-of-Experts (MoE)

Released

13.06.2026

License

Kimi K3 License

Open Weights Commercial Use Multimodal MXFP4/MXFP8 Kimi English Chinese

Input Modalities

text image video

Output Modalities

text

Context (native)

1,048,576 tokens

Context (extended)

1,048,576 tokens

Openness Index Score 70.0/100

About

Kimi K3 (moonshotai/Kimi-K3, released June 13, 2026 under the Kimi K3 License) is Moonshot AI's open-weight, native multimodal agentic model and its most capable model to date - a 2.8T-parameter MoE model (104B activated per token, 93 layers: 1 dense + 69 KDA + 24 Gated MLA attention layers; activates 16 of 896 experts), the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work and reasoning.

It is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and scales up MoE sparsity with a Stable LatentMoE framework yielding an approximate 2.5x improvement in overall scaling efficiency over Kimi K2. Kimi K3 understands text, images and video in the same model, with a 1M-token context window, a MoonViT-V2 vision encoder (401M parameters), SiTU-GLU activation, and native MXFP4 quantization (quantization-aware training from the SFT stage onward: MXFP4 weights / MXFP8 activations for broad hardware compatibility). It sustains long engineering sessions with minimal oversight - navigating massive repositories and orchestrating terminal tools from GPU kernels and compilers to vision-in-the-loop game dev, CAD and chip design - and advances end-to-end agentic knowledge work with interactive visualizations and deep research. Full weights are open for research, deployment and further innovation.

Training Data Quantization-aware training from SFT stage onward using MXFP4 weights with MXFP8 activations

Benchmark Scores

Benchmark Score Date
GPQA Diamond
stem_reasoning
98.49%
18.08.2026
CritPt (no tools)
stem_reasoning
71.92%
18.08.2026
AA-LCR
long_context
93.38%
18.08.2026
HLE (with tools)
stem_reasoning
81.99%
18.08.2026
DeepSWE 1.1
coding_agent
89.88%
18.08.2026
FrontierSWE
coding_agent
88.18%
18.08.2026
MLS-Bench-Lite
coding_agent
93.10%
18.08.2026
PostTrainBench
coding_agent
80.20%
18.08.2026
Terminal Bench 2.1
coding_agent
97.45%
18.08.2026
SciCode (subtask)
stem_reasoning
99.03%
18.08.2026
WildClawBench
coding_agent
73.53%
18.08.2026
BrowseComp
general_agent
100.00%
18.08.2026
DeepSearch QA
general_agent
100.00%
18.08.2026
MCP-Atlas
general_agent
97.95%
18.08.2026
MCPMark
general_agent
100.00%
18.08.2026
Automation-Bench
general_agent
45.45%
18.08.2026
JobBench
general_agent
70.35%
18.08.2026
Agents' Last Exam
general_agent
83.49%
18.08.2026
Apex-Agents
general_agent
97.79%
18.08.2026
OfficeQA Pro
general_agent
92.62%
18.08.2026
SpreadSheetBench-v1
general_agent
34.80
18.08.2026
Toolathlon Verified
general_agent
97.21%
18.08.2026
TauBench V3 Banking
general_agent
100.00%
18.08.2026
GDPVal-AA v2
general_agent
92.08%
18.08.2026
MMMU-Pro
vision_language
100.00%
18.08.2026
CharXiv (RQ)
document_understanding
100.00%
18.08.2026
MathVision
vision_language
100.00%
18.08.2026
ZEROBench
vision_language
82.61%
18.08.2026
MMVU
video_understanding
100.00%
18.08.2026
BabyVision
vision_language
94.94%
18.08.2026
VideoMME (w sub.)
video_understanding
100.00%
18.08.2026
Humanity's Last Exam
stem_reasoning
78.41%
13.08.2026
Cybergym
general_agent
83.60%
13.08.2026
DeepSWE
coding_agent
92.85%
13.08.2026
DSBench-FullStack
coding_agent
91.29%
13.08.2026
DSBench-Hard
coding_agent
81.05%
13.08.2026
Terminal-Bench 3.0
coding_agent
42.67%
28.08.2026
NL2Repo
coding_agent
73.38%
28.08.2026
ProgramBench
coding_agent
21.04%
28.08.2026
SWE-Marathon
coding_agent
96.12%
28.08.2026
ExploitGym (2h)
cybersecurity
10.89%
28.08.2026
ExploitGym (6h)
cybersecurity
16.48%
28.08.2026
ExploitBench
cybersecurity
14.55%
28.08.2026
SWE-bench Multilingual
coding_agent
89.70%
28.08.2026
SWE-bench Pro
coding_agent
79.12%
28.08.2026
SWE Atlas - QnA
coding_agent
47.45%
28.08.2026
SWE Atlas - TW
coding_agent
48.52%
28.08.2026
SWE Atlas - RF
coding_agent
59.43%
28.08.2026
Terminal-Bench 2.1 (Terminus-2)
coding_agent
99.26%
28.08.2026
Hy-Backend 2.0 (Internal)
coding_agent
50.00%
28.08.2026
Hy-SWE Max Verified (Internal)
coding_agent
78.67%
28.08.2026
Hy-CompanyBench V2 (Internal)
general_agent
78.09%
28.08.2026
WideSearch
general_agent
89.63%
28.08.2026
$OneMillion-Bench
general_capabilities
82.51%
28.08.2026
DRACO
general_agent
52.56%
28.08.2026
Hy-LifeSearch (Internal)
general_agent
28.16%
28.08.2026
Hy-BrowseComp-Pro2 (Internal)
general_agent
78.38%
28.08.2026
WorkSpaceBench
general_agent
40.48%
28.08.2026
BankerToolBench
general_agent
53.33%
28.08.2026
E-Bench (Internal)
general_agent
78.82%
28.08.2026
E-Bench-Code (Internal)
coding_agent
70.00%
28.08.2026
Hy-FinAgentBench (Internal)
domain_finance
66.67%
28.08.2026
Hy-FinmodelBench v2 (Internal)
domain_finance
63.64%
28.08.2026
BioMysteryBench
stem_reasoning
35.16%
28.08.2026
SUPERChem
stem_reasoning
59.34%
28.08.2026
ArXivMath
stem_reasoning
32.73%
28.08.2026
HorizonMath (pass@4)
stem_reasoning
50.28%
28.08.2026
MathArena Apex 2025
stem_reasoning
56.36%
28.08.2026
BrokenArXiv
stem_reasoning
58.04%
28.08.2026
SkillsBench Avg5
coding_agent
76.24%
28.08.2026
OSWorld-Verified
general_agent
100.00%
18.08.2026

Kimi K3 Evaluation Methodology

Evaluation Methodology

General Settings

  • All Kimi K3 results: reasoning_effort="max", temperature=1.0
  • Single-step tasks (GPQA Diamond, HLE-Full, vision benchmarks without tools): top-p=0.95
  • Agentic tasks: top-p=1.0
  • For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision, ZeroBench: each cell reports scores without and with tool augmentation (general tools for HLE-Full, Python for vision benchmarks)

Reasoning & Knowledge Benchmarks

Coding Benchmarks

  • DeepSWE: Kimi K3 evaluated with Kimi Code harness (67.3 with mini-SWE-agent harness); GLM-5.2 score from GLM-5.2 release blog; others from official DeepSWE leaderboard
  • Terminal-Bench 2.1: Kimi K3 with Kimi Code harness; other models report best score across harnesses
  • ProgramBench: Kimi K3 with Kimi Code harness; GLM-5.2 from release blog; others from Vals AI
  • SWE-Marathon: Kimi K3, Claude Opus 4.8, Claude Fable 5 with Claude Code harness; GPT-5.6 Sol with Codex harness; evaluation on H20-calibrated branch
  • FrontierSWE: Kimi K3 with Kimi Code harness; GPT-5.6 Sol with Codex harness; others from FrontierSWE
  • PostTrainBench: Kimi K3, Claude Fable 5 with Claude Code harness, GPT-5.6 Sol with Codex harness; averaged over 3 runs on H20 GPUs
  • MLS-Bench-Lite: Kimi K3 with Kimi Code harness; Claude models with Claude Code; GPT models with Codex
  • SciCode: Scores from Artificial Analysis (July 23, 2026)
  • Kimi Code Bench 2.0: In-house benchmark; Kimi K3 with Kimi Code harness (73.7 with Claude Code harness)

Agentic Benchmarks

  • BrowseComp: Context-compaction at 300K tokens; full 1M context scores 90.4
  • OfficeQA Pro: PDFs rendered as images, no machine-readable text
  • MCP-Atlas: 500-task public subset, 100-turn limit, Gemini 3.1 Pro as judge
  • AutomationBench: 600-task public subset, official GitHub setup
  • GDPval-AA v2, AA-Briefcase, τ³-Banking, Harvey Lab-AA, APEX-Agents: From Artificial Analysis and APEX-Agents leaderboard (July 23, 2026)
  • CorpFin v2, Finance Agent v2, Legal Research Bench: From Vals AI
  • Agents' Last Exam: From official leaderboard (July 23, 2026); each model paired with specific harness

Multimodal Benchmarks

  • Except ZeroBench (5 runs, official setting), all multimodal scores averaged over 3 runs
  • MMMU-Pro: official protocol, original input order, images prepended to text input
  • PerceptionBench: In-house benchmark for atomic visual perception

Kimi K3 Citation and Resources

Citation and Resources

Primary Sources

Resource URL
Tech Blog kimi.com/blog/kimi-k3
Full Technical Report (PDF) github.com/MoonshotAI/Kimi-K3/k3_tech_report.pdf
HuggingFace Model Page huggingface.co/moonshotai/Kimi-K3
API Platform platform.kimi.ai

Organization Links

Resource URL
Moonshot AI Website moonshot.ai
Kimi Product kimi.com
Moonshot AI on HuggingFace huggingface.co/moonshotai
Twitter twitter.com/kimi_moonshot
Discord discord.gg/TYU2fdJykW
ModelScope modelscope.cn/organization/moonshotai

Contact

Quickstart Guides

Guide URL
Kimi K3 Quickstart platform.kimi.ai/docs/guide/kimi-k3-quickstart
Thinking Effort Guide platform.kimi.ai/docs/guide/use-thinking-effort
Kimi Code CLI kimi.com/code

Kimi K3 Thinking Control and Reasoning Effort

Thinking Control

Always-On Thinking

Kimi K3 always has thinking enabled and returns reasoning_content in every response.

Reasoning Effort Levels

The reasoning_effort request field controls thinking depth:

Level Description
"low" Minimal reasoning, faster responses
"high" Moderate reasoning depth
"max" (default) Maximum reasoning depth

Preserved Thinking History

Kimi K3 was trained in preserved thinking history mode. This means:

  • For multi-turn conversations, the complete assistant message (including reasoning_content and tool_calls) must be passed back to messages as-is.
  • Only passing content (without reasoning_content) will degrade performance in multi-turn settings.
  • This is critical for tool calls — the thinking context must be preserved across turns.

Evaluation Settings

  • All benchmark results use reasoning_effort="max" with temperature=1.0
  • Single-step tasks: top-p=0.95
  • Agentic tasks: top-p=1.0

Additional Resources

Kimi K3 Training and Quantization Approach

Training Approach

Quantization-Aware Training (QAT)

Kimi K3 applies quantization-aware training from the SFT stage onward, using:

  • MXFP4 weights (4-bit microscaling float)
  • MXFP8 activations (8-bit microscaling float)

This means the model was trained with the quantization format in the loop, rather than being quantized post-hoc. This preserves model quality at reduced precision.

Architecture Training Stability

The Stable LatentMoE framework enables stable training of the high-sparsity MoE architecture (896 experts, 16 activated). This framework was key to achieving the 2.5× scaling efficiency improvement over Kimi K2.

Training Data

  • Languages: English, Chinese
  • Training data info: Quantization-aware training from SFT stage onward using MXFP4 weights with MXFP8 activations
  • Full details available in the Technical Report (PDF)

Kimi K3 Vision and Multimodal Capabilities

Vision and Multimodal Capabilities

Native Multimodality

Kimi K3 understands text, images, and video within the same model, powered by its native multimodal architecture.

Vision Encoder

  • Encoder: MoonViT-V2
  • Parameters: 401M
  • Supports text, image, and video inputs

Modalities

Direction Modalities
Input Text, Image, Video
Output Text

Vision Benchmark Performance

Benchmark Kimi K3 (max)
WorldVQA ForceAnswer 51.0
OmniDocBench 91.1
PerceptionBench 58.5
Video-MME (w. sub) 90.0
MMVU 82.1
BabyVision w/ python 85.7
MMMU-Pro 81.6 / 83.4
CharXiv (RQ) 84.8 / 91.3
MathVision 94.3 / 97.8
ZeroBench (pass@5) 23.0 / 41.0

Notes

  • Except for ZeroBench (5 runs, official setting), all multimodal scores are averaged over 3 runs.
  • MMMU-Pro is evaluated following the official protocol, preserving the original input order and prepending images to the text input.
  • PerceptionBench is an in-house benchmark focusing on atomic visual perception capabilities.

Kimi K3 Model Ecosystem

Model Ecosystem

Model Tree

Type Count Link
Adapters 1 model HF Models
Finetunes 43 models HF Models
Quantizations 47 models HF Models

Quantization Compatibility

The model page lists compatibility with:

  • llama.cpp
  • LM Studio
  • Jan
  • Ollama

Community

  • 171 discussions on HuggingFace
  • 46 Spaces using moonshotai/Kimi-K3
  • 2,787,971 downloads last month
  • 11 likes
  • 17,600 followers for Moonshot AI

Collection

Part of the Kimi K3 Collection — 2 items, 113 likes.

Kimi K3 Key Highlights

Key Highlights

Kimi K3 represents several breakthroughs as Moonshot AI's most capable model:

  1. World's first open 3T-class model — 2.8T parameters with open weights
  2. Novel KDA Architecture — Kimi Delta Attention + Attention Residuals, a departure from standard attention mechanisms
  3. 2.5× scaling efficiency over Kimi K2 via Stable LatentMoE framework (16/896 experts activated)
  4. Native multimodality — text, image, and video understanding in a single model
  5. 1M-token context window — native, no extension needed
  6. Long-horizon coding — operates with minimal human oversight across massive repositories
  7. Agentic knowledge work — deep research with interactive visualizations, dashboards, motion design
  8. Native MXFP4 quantization — quantization-aware training from SFT stage for broad hardware compatibility
  9. Open frontier weights — full weights under Kimi K3 License for research and commercial deployment

Kimi K3 Context Length

Context Length

Property Value
Native Context Length 1,048,576 tokens (1M)
Extended Context Length 1,048,576 tokens (1M)

Kimi K3 supports a 1-million-token context window natively. The model can process text, images, and video within this context.

Context Management

  • For BrowseComp evaluation, a context-compaction strategy is triggered at 300K tokens.
  • When evaluated with the full 1M-token context window and no context management, Kimi K3 achieves a BrowseComp score of 90.4 (vs. 91.2 with compaction).

Kimi K3 License Information

License

Both the code repository and the model weights are released under the Kimi K3 License.

  • The model is open-weight with commercial use allowed.
  • The license terms are available at the above URL.

Kimi K3 Coding Agent Framework

Coding Agent Framework

Kimi K3 works best with Kimi Code CLI as its agent framework. Users can run Kimi Code in their terminal and select Kimi K3 using the /model command.

Capabilities

  • Long-Horizon Coding: Sustains long engineering sessions with minimal human oversight
  • Repository Navigation: Navigates massive codebases effectively
  • Terminal Tool Orchestration: From GPU kernel optimization and compiler development to vision-in-the-loop game dev, CAD, and chip design
  • Multi-tool Coordination: Coordinates terminal tools, file editors, and build systems across complex workflows

Kimi K3 API Usage and Thinking Mode

Model Usage

Thinking Mode

Kimi K3 always has thinking enabled and will return reasoning_content. Thinking effort is configured with the top-level reasoning_effort request field:

  • "low" — minimal reasoning
  • "high" — moderate reasoning
  • "max" (default) — maximum reasoning

Preserved Thinking History Mode

Kimi K3 was trained in the preserved thinking history mode. For multi-turn conversations and tool calls, the complete assistant message returned by the API must be passed back to messages as-is — including reasoning_content and tool_calls, not just content.

Example: Multi-turn with Preserved Thinking

import openai

def chat_with_preserved_thinking(client: openai.OpenAI, model_name: str):
    messages = [
        {
            "role": "user",
            "content": "Tell me three random numbers."
        },
        {
            "role": "assistant",
            "reasoning_content": "I'll start by listing five numbers: 473, 921, 235, 215, 222, and I'll tell you the first three.",
            "content": "473, 921, 235"
        },
        {
            "role": "user",
            "content": "What are the other two numbers you have in mind?"
        }
    ]

    response = client.chat.completions.create(
        model=model_name,
        messages=messages,
        stream=False,
        max_tokens=4096,
        reasoning_effort="max",
    )
    # the assistant should mention 215 and 222 from prior reasoning content
    print(f"response: {response.choices[0].message.reasoning}")
    return response.choices[0].message.content

Additional Resources

Kimi K3 Deployment Options

Deployment

API Access

Kimi K3's API is available on https://platform.kimi.ai by selecting kimi-k3. The API provides OpenAI/Anthropic-compatible endpoints.

Supported Inference Engines

Engine Link Recipes
vLLM github.com/vllm-project/vllm recipes.vllm.ai/moonshotai/Kimi-K3
SGLang github.com/sgl-project/sglang docs.sglang.io/cookbook/.../Kimi-K3
TokenSpeed github.com/lightseekorg/tokenspeed lightseek.org/tokenspeed/recipes

Kimi K3 Native MXFP4 Quantization

Native MXFP4 Quantization

Kimi K3 applies quantization-aware training (QAT) from the SFT stage onward, using:

  • MXFP4 weights — 4-bit microscaling float format for model weights
  • MXFP8 activations — 8-bit microscaling float format for activations

This approach ensures broad hardware compatibility while maintaining model quality. The quantization is native (baked into the model during training), not a post-hoc compression step, which preserves the model's reasoning and coding capabilities at the reduced precision.

Tensor Types in Released Weights

  • F32
  • BF16
  • U8

The model is distributed in Safetensors format with the native MXFP4 quantization already applied.

Kimi K3 Benchmark Results

Evaluation Results

All Kimi K3 results obtained with reasoning effort set to 'max' and temperature = 1.0. For single-step tasks (GPQA Diamond, HLE-Full, vision benchmarks without tools), top-p = 0.95; for agentic tasks, top-p = 1.0.

Reasoning & Knowledge

Benchmark Kimi K3 (max)
GPQA Diamond 93.5
CritPt 23.4
AA-LCR 74.7
HLE-Full 43.5 / 56.0

Coding

Benchmark Kimi K3 (max)
DeepSWE 67.5
ProgramBench 77.8
Terminal-Bench 2.1 88.3
FrontierSWE 81.2
SWE-Marathon 42.0
PostTrainBench 36.6
MLS-Bench-Lite 48.3
SciCode 58.7
Kimi Code Bench 2.0 72.9

Agentic

Benchmark Kimi K3 (max)
BrowseComp 91.2
DeepSearchQA (F1) 95.0
ResearchRubrics 76.2
GDPval-AA v2 (Elo) 1686
Toolathlon-Verified 76.5
MCPMark-Verified 94.5
MCP-Atlas 84.2
AutomationBench 30.8
JobBench 54.3
AA-Briefcase (Elo) 1548
Agents' Last Exam 28.3
APEX-Agents 41.0
OfficeQA Pro 63.3
SpreadsheetBench 2 34.8
OSWorld-Verified 84.8
OSWorld 2.0 58.3
SaaS-Bench 60.1
τ³-Banking 33.4
Harvey Lab-AA 94.6
CorpFin v2 71.6
Finance Agent v2 54.4
Legal Research Bench 44.2

Vision

Benchmark Kimi K3 (max)
WorldVQA ForceAnswer 51.0
OmniDocBench 91.1
PerceptionBench 58.5
Video-MME (w. sub) 90.0
MMVU 82.1
BabyVision w/ python 85.7
MMMU-Pro 81.6 / 83.4
CharXiv (RQ) 84.8 / 91.3
MathVision 94.3 / 97.8
ZeroBench (pass@5) 23.0 / 41.0

Footnotes

  • CritPt and AA-LCR scores cited from Artificial Analysis (July 23, 2026).
  • For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision, and ZeroBench, each cell reports scores without and with tool augmentation.
  • DeepSWE: Kimi K3 evaluated with Kimi Code harness (67.3 with mini-SWE-agent harness).
  • BrowseComp: Context-compaction strategy triggered at 300K tokens; full 1M-token context scores 90.4.

Kimi K3 Architecture Details

Architecture Summary

Property Value
Architecture Mixture-of-Experts (MoE)
Total Parameters 2.8T
Activated Parameters 104B
Number of Layers 93
Number of Dense Layers 1
Attention-Layer Composition 69 KDA + 24 Gated MLA
Attention Hidden Dimension 7168
Number of Attention Heads 96
Latent MoE Dimension 3584
MoE Hidden Dimension (per Expert) 3072
Number of Experts 896
Selected Experts per Token 16
Number of Shared Experts 2
Vocabulary Size 160K
Context Length 1,048,576
Attention Mechanism KDA & Gated MLA
Activation Function SiTU-GLU
Vision Encoder MoonViT-V2
Parameters of Vision Encoder 401M
Quantization MXFP4 weights / MXFP8 activations (quantization-aware training)
Modality Text, Image

Key Architectural Innovations

  • Kimi Delta Attention (KDA): Novel attention mechanism used in 69 of 93 layers.
  • Gated MLA: Used in 24 layers, complementing KDA for a hybrid attention design.
  • Attention Residuals (AttnRes): Residual connections across attention layers for improved gradient flow.
  • Stable LatentMoE: Framework for stable training of high-sparsity MoE models.
  • SiTU-GLU: Custom activation function combining SiLU with a Gated Linear Unit.

Kimi K3 Model Introduction

Model Introduction

Kimi K3 is an open-weight, native multimodal agentic model and Moonshot AI's most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

Key Features

  • New Architecture: Built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), scaling up MoE sparsity with a Stable LatentMoE framework that activates 16 out of 896 experts — yielding an approximate 2.5× improvement in overall scaling efficiency over Kimi K2.
  • Long-Horizon Coding: Operates with minimal human oversight, sustains long engineering sessions, navigates massive repositories, and orchestrates terminal tools — from GPU kernel optimization and compiler development to vision-in-the-loop game dev, CAD, and chip design.
  • Agentic Knowledge Work: Produces deep research with interactive visualizations, widgets and dashboards, and motion design and video editing, powered by its native multimodal architecture.
  • Native Multimodality & Long Context: Understands text, images, and video within the same model; supports a 1-million-token context window.
  • Open Frontier Weights: Full model weights released under the Kimi K3 License for research, deployment, and further innovation.

Architecture

Decoder Block input Embedding vocab 164K · d 7168 Full Attention MLA · 96 heads ×93 MoE FFN 896 experts · +2 shared · dᴻ 3072 Final Norm LM Head vocab 164K output
Attention
Multi-head Latent Attention
MoE
896 experts
Layers
93
Hidden size
7168
Context
1M tokens
Parameters
2800000M
Active params
104000M

Source: Hugging Face config.json · KimiK3ForConditionalGeneration · model repo

Type: Mixture-of-Experts (MoE)
Attention: KDA & Gated MLA
Decoder: Kimi Delta Attention (KDA) + Attention Residuals (AttnRes)
MoE: yes (896 experts)
Routing: Stable LatentMoE - activates 16 out of 896 experts with 2 shared experts
Layers 93
Context length 1M
Extended context 1M
Experts 896
Experts per token 16
Shared experts 2
Attention heads 96
Hidden size 7168
Attention hidden size 7168
Vocabulary 160K
Expert FFN dim 3072
Activation SiTU-GLU
Vision Yes
Vision encoder MoonViT-V2
Dense Layers 1
Gated Mla Layers 24
Kda Layers 69
Latent Moe Dim 3584
MTP Yes
Quantization MXFP4 weights / MXFP8 activations (quantization-aware training)
Vision encoder params 401M

Training Pipeline

  1. 1
    pretraining

    Pretraining

    2.8T parameter MoE model with Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). Uses Stable LatentMoE framework activating 16 out of 896 experts. Native MXFP4 quantization-aware training from SFT stage onward.

  2. 2
    sft

    Supervised Fine-Tuning (SFT)

    Quantization-aware training applied from SFT stage onward using MXFP4 weights with MXFP8 activations for broad hardware compatibility.

  3. 3
    rl

    Reinforcement Learning (RL)

    Trained in preserved thinking history mode. Thinking effort supports low, high, and max settings.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
DeepSWE evaluation — —
GPQA Diamond evaluation — —
WildClawBench evaluation — —
MDPBench evaluation — —

Trend Analysis

24h Change

+0.2%

7d Change

+2.0%

Current

18,018

huggingface

downloads_all_time

+3.5%

huggingface

likes

+0.1%

ollama

downloads

+2.1%

huggingface

downloads

-0.3%

View raw metric history →

Usage & Social Metrics

SourceMetricValuePeriodRecorded
ollama downloads 68,700 pulls daily 01.09.2026
huggingface downloads_all_time 3,735,508 daily 01.09.2026
huggingface followers 18,018 daily 01.09.2026
huggingface likes 11,128 daily 01.09.2026
huggingface downloads 2,783,061 daily 01.09.2026
ollama downloads 67,300 pulls daily 31.08.2026
huggingface downloads_all_time 3,610,900 daily 31.08.2026
huggingface followers 17,977 daily 31.08.2026
huggingface likes 11,116 daily 31.08.2026
huggingface downloads 2,792,274 daily 31.08.2026
ollama downloads 65,700 pulls daily 30.08.2026
huggingface downloads_all_time 3,446,843 daily 30.08.2026
huggingface followers 17,928 daily 30.08.2026
huggingface likes 11,099 daily 30.08.2026
huggingface downloads 2,794,721 daily 30.08.2026
ollama downloads 64,200 pulls daily 29.08.2026
huggingface followers 17,884 daily 29.08.2026
huggingface likes 11,080 daily 29.08.2026
huggingface downloads 2,701,014 daily 29.08.2026
ollama downloads 62,000 pulls daily 28.08.2026
huggingface followers 17,836 daily 28.08.2026
huggingface likes 11,061 daily 28.08.2026
huggingface downloads 2,675,145 daily 28.08.2026
ollama downloads 60,300 pulls daily 27.08.2026
huggingface followers 17,787 daily 27.08.2026
huggingface likes 11,036 daily 27.08.2026
huggingface downloads 2,829,554 daily 27.08.2026
ollama downloads 58,200 pulls daily 26.08.2026
huggingface followers 17,709 daily 26.08.2026
huggingface likes 11,018 daily 26.08.2026
huggingface downloads 2,921,257 daily 26.08.2026
ollama downloads 55,600 pulls daily 25.08.2026
huggingface followers 17,658 daily 25.08.2026
huggingface likes 10,995 daily 25.08.2026
huggingface downloads 2,865,293 daily 25.08.2026
ollama downloads 53,000 pulls daily 24.08.2026
huggingface followers 17,609 daily 24.08.2026
huggingface likes 10,972 daily 24.08.2026
huggingface downloads 2,787,971 daily 24.08.2026
huggingface followers 17,556 daily 23.08.2026
huggingface likes 10,949 daily 23.08.2026
huggingface downloads 2,727,920 daily 23.08.2026
huggingface followers 17,515 daily 22.08.2026
huggingface likes 10,924 daily 22.08.2026
huggingface downloads 2,612,739 daily 22.08.2026
huggingface followers 17,482 daily 21.08.2026
huggingface likes 10,913 daily 21.08.2026
huggingface downloads 2,448,810 daily 21.08.2026
huggingface followers 17,418 daily 20.08.2026
huggingface likes 10,883 daily 20.08.2026

View full metric history →

Related Models