DeepSeek is the open-weight model family that changed how the
industry talks about training cost. Its January 2025 release,
R1, didn't just perform well, it disclosed roughly how much
compute went into the base model it was built on, and that number
was small enough to wipe out hundreds of billions of dollars in AI
infrastructure stock value in a single day. Here's who builds it,
how it's actually trained, and when it's the right call.
Who builds DeepSeek
DeepSeek AI, based in Hangzhou, was founded in
July 2023 by Liang Wenfeng, who also founded the
quant hedge fund High-Flyer, which had used
AI-driven trading since 2021 and largely bankrolled DeepSeek's
compute and research. Unusually for a frontier lab, DeepSeek took
no outside venture capital for its first three years, until June
2026, when it closed roughly $7.4 billion led by Tencent and CATL
at a reported $50-59 billion valuation, with Liang retaining
majority ownership. Liang has also spoken publicly about running
the company without mandatory overtime or individual KPI tracking,
a deliberate contrast to China's "996" work culture.
DeepSeek's defining moment came on January 20, 2025, with the
release of R1, a reasoning model that appeared
competitive with frontier reasoning systems while DeepSeek
disclosed a training run for its underlying V3 base model of about
2.79 million H800 GPU-hours, far below what the market assumed
frontier-level training required. The resulting sell-off wiped
roughly $589 billion off Nvidia's market capitalization in a
single day, the largest single-day loss in stock market history at
the time, with Broadcom and ASML falling sharply as well. DeepSeek
has also drawn export-control scrutiny, including allegations
reported by Bloomberg and a US House committee about its use of
restricted Nvidia chips, allegations DeepSeek has not
substantively addressed publicly and that remain contested rather
than settled. Release-wise, the family has moved from V2 (May
2024) through V3 and R1 (late 2024/January 2025), a hybrid
thinking/non-thinking V3.1 (August 2025), V3.2 with a new sparse
attention mechanism (September-December 2025), to V4 (April 2026).
A rumored R2 has not been officially released as of this writing;
treat any specs you see for it as unconfirmed.
V3 and R1's January 2025 release is the inflection point the rest of the industry still references.
How it's trained
DeepSeek-V3, per its technical report, is a
671B-parameter mixture-of-experts model with 37B
active per token, built on Multi-head Latent Attention,
which compresses the key-value cache for cheaper inference, and
DeepSeekMoE, a fine-grained expert design with
shared experts and a load-balancing scheme that avoids the
auxiliary loss term most MoE models use to keep experts evenly
utilized, DeepSeek balances load without one. Training also used a
multi-token prediction objective, predicting
several future tokens per step rather than just the next one, and
FP8 mixed-precision training, one of the first
successful demonstrations of FP8 at this scale. DeepSeek pretrained
V3 on 14.8 trillion tokens with a 128K context window, and
disclosed the run's compute cost directly: about 2.79 million
H800 GPU-hours. The widely repeated "$5.5-6 million" price tag
is an analyst estimate applied to that GPU-hour figure, not a
dollar amount DeepSeek itself published, worth keeping straight
since the estimate is what triggered the market reaction, not an
official number.
R1's training is arguably the more novel part.
R1-Zero, an intermediate model, was trained with
pure large-scale reinforcement learning directly on the V3 base
checkpoint, no supervised fine-tuning warm-start at all, rewarding
only final-answer correctness and output format. That process
produced genuinely emergent behavior, the model learned to pause
and re-check its own reasoning without being explicitly trained to
do so, but also had readability problems and mixed languages
within a single response. R1 itself fixes that by adding a small
supervised "cold-start" dataset before RL, plus further SFT and RL
stages on top. The RL algorithm behind both, Group
Relative Policy Optimization (GRPO), originated in
DeepSeek's earlier DeepSeekMath work: it drops the separate critic
model that standard PPO needs and instead computes each output's
advantage relative to a group of other sampled outputs for the
same prompt, which is a meaningful part of why DeepSeek could run
this at scale without a second model's worth of extra memory and
compute. DeepSeek then distilled roughly 800,000 reasoning traces
generated by R1 into six smaller dense models, Qwen-based at
1.5B/7B/14B/32B and Llama-based at 8B/70B, using supervised
fine-tuning alone with no RL, released as the R1-Distill line.
Later releases kept iterating on the same base: V3.1 (August
2025) unified thinking and non-thinking modes into one model
instead of splitting them across separate endpoints, and V3.2
(September-December 2025) added DeepSeek Sparse Attention for
cheaper long-context inference alongside further agentic
tool-use RL. V4 (April 2026) ships as two sizes, V4-Pro at 1.6
trillion total parameters with 49B active, and V4-Flash at 284B
total with 13B active, both with a 1-million-token context window
and reportedly pretrained on over 32 trillion tokens. V4's deeper
architectural changes are documented mainly through Hugging Face
model cards and third-party technical writeups rather than a
peer-reviewed report the way V3 and R1 were, so treat its exact
internals as reported rather than fully verified.
TRAINING EFFICIENCY
FP8 training, auxiliary-loss-free load balancing, and multi-token prediction are the specific techniques behind DeepSeek's disclosed compute figures, not just a smaller model.
MATH & REASONING
R1's pure-RL training pipeline and its successors have posted gold-medal-level results on competitions like the IMO, CMO, and ICPC World Finals.
DISTILLED FOR SELF-HOSTING
The R1-Distill dense models, from 1.5B to 70B, put R1-style reasoning behavior on ordinary Qwen and Llama-sized checkpoints that fit on a single GPU.
When to use it
DeepSeek's core pitch is frontier-comparable performance,
especially on math, coding, and multi-step reasoning, at a
fraction of the training and serving cost most labs assumed was
required. By 2026, it sits alongside Kimi, Qwen, and Zhipu's GLM
as one of several leading Chinese open-weight families rather than
a singular outlier, each with a slightly different edge: DeepSeek
is generally cited as strong on agentic coding and graduate-level
reasoning, Kimi on long-horizon multi-step agent tasks, GLM on
cost-efficient day-to-day coding, and Qwen on breadth and adoption.
Against closed Western frontier models, DeepSeek's own
benchmarks position V3.2 and V4 as competitive on reasoning at a
much lower API price, though those comparisons come from
DeepSeek's own marketing and are worth verifying independently
rather than taken at face value.
In practice, the size decision matters more than the model
choice. The full V3 or V4 flagship models, at 671B to 1.6 trillion
total parameters, need real multi-node infrastructure to self-host,
not a single card. If you want R1-style reasoning behavior without
that infrastructure, the R1-Distill dense models at 1.5B through
70B are the practical target, they trade some of R1's ceiling for
running on hardware you likely already have.
Reasoning behavior distilled into existing dense architectures, not a shrunk version of R1's own MoE design.
How to use it
DeepSeek's model weights, including V3, R1, the R1-Distill line,
and V4, are released under a permissive MIT license,
allowing commercial use, modification, and self-hosting without
royalty, one of the most straightforward licenses of any frontier-
adjacent model family. Weights are published on Hugging Face under
the deepseek-ai org, with inference code and technical
reports on GitHub. For serving, vLLM and
SGLang both typically ship support close to
release day, SGLang in particular is notable for AMD GPU
compatibility in both FP8 and BF16, and TensorRT-LLM
is commonly supported for NVIDIA-optimized deployment. The full
MoE flagship models require a real multi-GPU cluster to self-host
given their total parameter counts; the R1-Distill dense models are
the models actually meant for single or few-GPU self-hosting.
Beyond self-hosting, DeepSeek runs its own API platform, widely
noted for aggressive pricing relative to comparable closed models,
though exact rates have changed more than once and are worth
checking directly against DeepSeek's own pricing page rather than
a cached figure before budgeting around it. Fine-tuning support is
broad through standard Hugging Face transformers
tooling plus common community frameworks like Unsloth and Axolotl,
with the distilled dense models again being the practical target
for teams fine-tuning on their own data rather than the full-scale
MoE checkpoints.
Where this fits at Numerata
Whether you're working with a full DeepSeek flagship or one of
the R1-Distill models, P95 is where
you fine-tune it on your own data, on infrastructure you control,
private cloud or fully air-gapped, and NinetyFive
serves it afterward, so a model built around efficient training
also gets efficient serving, rather than losing that advantage to
a naive deployment.