BLOG · SLM SERIES

Kimi Explained

Kimi is the open-weight model family built by Moonshot AI, a Beijing lab founded in 2023. Its flagship line, Kimi K2, made its name in mid-2025 as a trillion-parameter model trained with a novel optimizer and pointed squarely at agentic tool use rather than chat, and by mid-2026 its successor K3 had pushed that further into what Moonshot calls the largest open-weight model released to date. Here's who builds it, how it's actually trained, and when it makes sense to reach for it.

Who builds Kimi

Moonshot AI (月之暗面, "dark side of the moon") was founded in Beijing in March 2023 by Yang Zhilin, Zhou Xinyu, and Wu Yuxin, former Tsinghua classmates; Yang previously worked at Carnegie Mellon and Meta FAIR. The company has raised roughly $3.77 billion from investors including Alibaba, Tencent, and Meituan, and was reported at around a $20 billion valuation as of mid-2026, with speculation of a Hong Kong IPO. Kimi K2's July 2025 release was widely described in industry commentary as another "DeepSeek moment," evidence that a Chinese lab could ship an open-weight model competitive with closed frontier systems on agentic and tool-use tasks specifically, not just benchmarks in general. Releases since have been fast: K2 (July 2025), a K2 Thinking variant with interleaved reasoning and tool calls (November 2025), K2.5 (January 2026), K2.6 (April 2026), and K3 (July 2026).

Kimi release timeline A timeline showing Kimi K2's release in July 2025, its Thinking variant that November, K2.5 and K2.6 through early 2026, and K3 in July 2026 as the current flagship. K2 JULY 2025 K2 THINKING NOV 2025 K2.5 / K2.6 EARLY 2026 K3 JULY 2026
Roughly two major releases a year since K2's debut, each pushed further into agentic tool use.

How it's trained

Kimi K2, per Moonshot's technical report, is a sparse mixture-of-experts model: 1.04 trillion total parameters, of which only about 32 billion activate per token, across 61 transformer layers with 384 experts (8 routed plus 1 shared per token) and multi-head latent attention. It pretrained on 15.5 trillion tokens, and Moonshot reports the run completed with zero loss spikes, a notable stability claim at trillion-parameter scale. That stability is credited to MuonClip, Moonshot's own optimizer, which combines the token-efficient Muon optimizer with a mechanism called QK-Clip that keeps attention logits from blowing up during training. It's the family's signature infrastructure contribution, and the main reason Kimi K2 gets cited in training-stability discussions independent of its benchmark scores.

The other defining piece of K2's training is what happens after pretraining. Rather than optimizing post-training mainly for math or code correctness, Moonshot built a large-scale synthetic data pipeline that generates simulated tools, agent personas, tasks, and multi-step trajectories, then verifies which trajectories actually succeed, to produce training data for tool use and long-horizon agency specifically. That's why K2 is described in its own paper's title as "open agentic intelligence" rather than a general chat model that happens to support tools. The November 2025 K2 Thinking variant pushed this further with interleaved reasoning and tool calls, meaning the model can invoke a tool mid-chain-of-thought instead of finishing a reasoning pass before acting, and reportedly sustains 200-300 sequential tool calls autonomously.

K3, released July 2026, is a real architectural departure rather than a scale-up of the same design. It's a 2.8 trillion parameter mixture-of-experts model with 104 billion parameters active per token, across 93 layers combining 69 Kimi Delta Attention (a hybrid linear attention mechanism Moonshot built in-house) layers with 24 gated multi-head latent attention layers, and routes each token through 16 of 896 experts plus 2 shared experts, which Moonshot says delivers roughly 2.5x better scaling efficiency than K2's architecture. It's natively multimodal (text, image, and video), supports a 1-million-token context window, and runs with an always-on "thinking mode" with selectable low, high, or max reasoning effort, rather than K2's separate thinking/non-thinking checkpoints.

AGENTIC TOOL USE

Trained on synthetic multi-step tool-use trajectories rather than chat data alone, K2 and its successors are purpose-built for long-horizon agent workflows, not adapted for them after the fact.

CODING

Strong on coding benchmarks like SWE-bench Verified and LiveCodeBench, in the same competitive tier as DeepSeek-V3 and Qwen's coder variants among open-weight models.

TRAINING STABILITY

The MuonClip optimizer is what let Moonshot train a trillion-parameter-plus MoE with zero reported loss spikes, an infrastructure result that gets cited well beyond Kimi's own benchmark scores.

When to use it

Kimi K2's own framing is the best guide to when it fits: reach for it when the task is genuinely agentic, a workflow that calls tools repeatedly, keeps state across many steps, and needs to recover from intermediate failures, rather than a single-turn question-answering task a smaller or cheaper model would handle fine. Its coding and tool-use benchmarks put it in the same competitive band as DeepSeek-V3 and Qwen's larger coder-focused models, and by mid-2026 K3 specifically is being discussed alongside DeepSeek and Qwen's own frontier tiers as one of the leading open-weight options overall. What it's not, is a small or cheap-to-serve model: even K2's 32B active parameters, let alone K3's 104B, means self-hosting either requires serious multi-GPU infrastructure, not a single card.

If your task doesn't actually need long-horizon tool use, agent memory across many steps, or frontier-tier coding, a smaller dense or lower-active-parameter MoE model, from Qwen's 8B-32B range or a similarly sized alternative, will usually serve faster and cheaper for the same result. Kimi earns its cost specifically on the agentic end of the spectrum.

Kimi K2 versus K3 scale and architecture A comparison showing Kimi K2 at 1.04 trillion total parameters with 32 billion active per token and a 128,000 token context window, versus Kimi K3 at 2.8 trillion total parameters with 104 billion active per token and a 1 million token context window, alongside K3's added native multimodal support. K2 1.04T TOTAL 32B ACTIVE 128K CONTEXT K3 2.8T TOTAL 104B ACTIVE NATIVE MULTIMODAL 1M CONTEXT
K3 roughly triples K2's active-parameter count and adds native multimodality and 8x the context window.

How to use it

Kimi K2, K2 Thinking, K2.5, and K2.6 all ship under a modified MIT license: standard MIT terms, plus one added condition, if a product built on it reaches more than 100 million monthly active users or $20 million in monthly revenue, it must visibly credit "Kimi K2" (or the specific version) in its user interface. Below that threshold, it behaves like ordinary permissive open source. K3 ships under its own separate license rather than reusing K2's terms verbatim, so check the license file in K3's Hugging Face repo before planning a commercial deployment around it rather than assuming the K2 terms carry over unchanged.

Weights for the whole family are published on Hugging Face under the moonshotai org, with deployment guides in the accompanying GitHub repos. Officially supported serving engines are vLLM, SGLang, KTransformers, and TensorRT-LLM; K2 ships in block-fp8, while K3 uses MXFP4 weights with MXFP8 activations via quantization-aware training, and vLLM/SGLang both expose Kimi-specific tool-call and reasoning parser flags for serving it correctly. Community GGUF quantizations (via Unsloth, among others) exist for local experimentation, though a model at this scale still needs substantial memory even quantized. Beyond self-hosting, Moonshot's own platform.moonshot.ai offers an OpenAI/Anthropic-compatible API, and Kimi.com and its mobile app offer a hosted chat interface if you want to evaluate the model before committing to serving it yourself.

Where this fits at Numerata

A model this size is exactly the case where control over your own serving stack matters most: P95 is where you'd fine-tune a Kimi checkpoint on your own agentic trajectories or task data, on infrastructure you control, and NinetyFive is what serves it afterward, so a model with this many active parameters still hits production latency instead of the multi-second response times naive multi-GPU serving tends to produce at this scale.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog