Kimi is the open-weight model family built by Moonshot AI, a
Beijing lab founded in 2023. Its flagship line, Kimi K2, made its
name in mid-2025 as a trillion-parameter model trained with a
novel optimizer and pointed squarely at agentic tool use rather
than chat, and by mid-2026 its successor K3 had pushed that further
into what Moonshot calls the largest open-weight model released to
date. Here's who builds it, how it's actually trained, and when it
makes sense to reach for it.
Who builds Kimi
Moonshot AI (月之暗面, "dark side of the moon")
was founded in Beijing in March 2023 by Yang Zhilin, Zhou Xinyu, and
Wu Yuxin, former Tsinghua classmates; Yang previously worked at
Carnegie Mellon and Meta FAIR. The company has raised roughly $3.77
billion from investors including Alibaba, Tencent, and Meituan, and
was reported at around a $20 billion valuation as of mid-2026, with
speculation of a Hong Kong IPO. Kimi K2's July 2025 release was
widely described in industry commentary as another "DeepSeek
moment," evidence that a Chinese lab could ship an open-weight
model competitive with closed frontier systems on agentic and
tool-use tasks specifically, not just benchmarks in general.
Releases since have been fast: K2 (July 2025), a K2 Thinking variant
with interleaved reasoning and tool calls (November 2025), K2.5
(January 2026), K2.6 (April 2026), and K3 (July 2026).
Roughly two major releases a year since K2's debut, each pushed further into agentic tool use.
How it's trained
Kimi K2, per Moonshot's technical report, is a sparse
mixture-of-experts model: 1.04 trillion total
parameters, of which only about 32 billion activate per token,
across 61 transformer layers with 384 experts (8 routed plus 1
shared per token) and multi-head latent attention. It pretrained on
15.5 trillion tokens, and Moonshot reports the run completed with
zero loss spikes, a notable stability claim at
trillion-parameter scale. That stability is credited to
MuonClip, Moonshot's own optimizer, which combines
the token-efficient Muon optimizer with a mechanism called QK-Clip
that keeps attention logits from blowing up during training. It's
the family's signature infrastructure contribution, and the main
reason Kimi K2 gets cited in training-stability discussions
independent of its benchmark scores.
The other defining piece of K2's training is what happens after
pretraining. Rather than optimizing post-training mainly for math
or code correctness, Moonshot built a large-scale synthetic data
pipeline that generates simulated tools, agent personas, tasks, and
multi-step trajectories, then verifies which trajectories actually
succeed, to produce training data for tool use and long-horizon
agency specifically. That's why K2 is described in its own paper's
title as "open agentic intelligence" rather than a general chat
model that happens to support tools. The November 2025 K2 Thinking
variant pushed this further with interleaved reasoning and tool
calls, meaning the model can invoke a tool mid-chain-of-thought
instead of finishing a reasoning pass before acting, and reportedly
sustains 200-300 sequential tool calls autonomously.
K3, released July 2026, is a real architectural departure rather
than a scale-up of the same design. It's a 2.8 trillion parameter
mixture-of-experts model with 104 billion parameters active per
token, across 93 layers combining 69 Kimi Delta Attention (a hybrid
linear attention mechanism Moonshot built in-house) layers with 24
gated multi-head latent attention layers, and routes each token
through 16 of 896 experts plus 2 shared experts, which Moonshot
says delivers roughly 2.5x better scaling efficiency than K2's
architecture. It's natively multimodal (text, image, and video),
supports a 1-million-token context window, and runs with an
always-on "thinking mode" with selectable low, high, or max
reasoning effort, rather than K2's separate thinking/non-thinking
checkpoints.
AGENTIC TOOL USE
Trained on synthetic multi-step tool-use trajectories rather than chat data alone, K2 and its successors are purpose-built for long-horizon agent workflows, not adapted for them after the fact.
CODING
Strong on coding benchmarks like SWE-bench Verified and LiveCodeBench, in the same competitive tier as DeepSeek-V3 and Qwen's coder variants among open-weight models.
TRAINING STABILITY
The MuonClip optimizer is what let Moonshot train a trillion-parameter-plus MoE with zero reported loss spikes, an infrastructure result that gets cited well beyond Kimi's own benchmark scores.
When to use it
Kimi K2's own framing is the best guide to when it fits: reach
for it when the task is genuinely agentic, a workflow that calls
tools repeatedly, keeps state across many steps, and needs to
recover from intermediate failures, rather than a single-turn
question-answering task a smaller or cheaper model would handle
fine. Its coding and tool-use benchmarks put it in the same
competitive band as DeepSeek-V3 and Qwen's larger coder-focused
models, and by mid-2026 K3 specifically is being discussed alongside
DeepSeek and Qwen's own frontier tiers as one of the leading
open-weight options overall. What it's not, is a small or
cheap-to-serve model: even K2's 32B active parameters, let alone
K3's 104B, means self-hosting either requires serious multi-GPU
infrastructure, not a single card.
If your task doesn't actually need long-horizon tool use, agent
memory across many steps, or frontier-tier coding, a smaller
dense or lower-active-parameter MoE model, from Qwen's 8B-32B
range or a similarly sized alternative, will usually serve faster
and cheaper for the same result. Kimi earns its cost specifically
on the agentic end of the spectrum.
K3 roughly triples K2's active-parameter count and adds native multimodality and 8x the context window.
How to use it
Kimi K2, K2 Thinking, K2.5, and K2.6 all ship under a
modified MIT license: standard MIT terms, plus one
added condition, if a product built on it reaches more than 100
million monthly active users or $20 million in monthly revenue, it
must visibly credit "Kimi K2" (or the specific version) in its
user interface. Below that threshold, it behaves like ordinary
permissive open source. K3 ships under its own separate license
rather than reusing K2's terms verbatim, so check the license file
in K3's Hugging Face repo before planning a commercial deployment
around it rather than assuming the K2 terms carry over unchanged.
Weights for the whole family are published on Hugging Face under
the moonshotai org, with deployment guides in the
accompanying GitHub repos. Officially supported serving engines are
vLLM, SGLang, KTransformers, and TensorRT-LLM; K2
ships in block-fp8, while K3 uses MXFP4 weights with MXFP8
activations via quantization-aware training, and vLLM/SGLang both
expose Kimi-specific tool-call and reasoning parser flags for
serving it correctly. Community GGUF quantizations (via Unsloth,
among others) exist for local experimentation, though a model at
this scale still needs substantial memory even quantized. Beyond
self-hosting, Moonshot's own platform.moonshot.ai offers an
OpenAI/Anthropic-compatible API, and Kimi.com and its mobile app
offer a hosted chat interface if you want to evaluate the model
before committing to serving it yourself.
Where this fits at Numerata
A model this size is exactly the case where control over your
own serving stack matters most: P95 is
where you'd fine-tune a Kimi checkpoint on your own agentic
trajectories or task data, on infrastructure you control, and
NinetyFive is what serves it
afterward, so a model with this many active parameters still hits
production latency instead of the multi-second response times
naive multi-GPU serving tends to produce at this scale.