BLOG · SLM SERIES

DeepSeek Explained

DeepSeek is the open-weight model family that changed how the industry talks about training cost. Its January 2025 release, R1, didn't just perform well, it disclosed roughly how much compute went into the base model it was built on, and that number was small enough to wipe out hundreds of billions of dollars in AI infrastructure stock value in a single day. Here's who builds it, how it's actually trained, and when it's the right call.

Who builds DeepSeek

DeepSeek AI, based in Hangzhou, was founded in July 2023 by Liang Wenfeng, who also founded the quant hedge fund High-Flyer, which had used AI-driven trading since 2021 and largely bankrolled DeepSeek's compute and research. Unusually for a frontier lab, DeepSeek took no outside venture capital for its first three years, until June 2026, when it closed roughly $7.4 billion led by Tencent and CATL at a reported $50-59 billion valuation, with Liang retaining majority ownership. Liang has also spoken publicly about running the company without mandatory overtime or individual KPI tracking, a deliberate contrast to China's "996" work culture.

DeepSeek's defining moment came on January 20, 2025, with the release of R1, a reasoning model that appeared competitive with frontier reasoning systems while DeepSeek disclosed a training run for its underlying V3 base model of about 2.79 million H800 GPU-hours, far below what the market assumed frontier-level training required. The resulting sell-off wiped roughly $589 billion off Nvidia's market capitalization in a single day, the largest single-day loss in stock market history at the time, with Broadcom and ASML falling sharply as well. DeepSeek has also drawn export-control scrutiny, including allegations reported by Bloomberg and a US House committee about its use of restricted Nvidia chips, allegations DeepSeek has not substantively addressed publicly and that remain contested rather than settled. Release-wise, the family has moved from V2 (May 2024) through V3 and R1 (late 2024/January 2025), a hybrid thinking/non-thinking V3.1 (August 2025), V3.2 with a new sparse attention mechanism (September-December 2025), to V4 (April 2026). A rumored R2 has not been officially released as of this writing; treat any specs you see for it as unconfirmed.

DeepSeek release timeline A timeline from DeepSeek-V2 in May 2024 through V3 and R1 in late 2024 and January 2025, the hybrid-reasoning V3.1 and sparse-attention V3.2 through 2025, to V4 in April 2026 as the current generation. V2 MAY 2024 V3 / R1 DEC 2024 / JAN 2025 V3.1 / V3.2 2025 V4 APRIL 2026
V3 and R1's January 2025 release is the inflection point the rest of the industry still references.

How it's trained

DeepSeek-V3, per its technical report, is a 671B-parameter mixture-of-experts model with 37B active per token, built on Multi-head Latent Attention, which compresses the key-value cache for cheaper inference, and DeepSeekMoE, a fine-grained expert design with shared experts and a load-balancing scheme that avoids the auxiliary loss term most MoE models use to keep experts evenly utilized, DeepSeek balances load without one. Training also used a multi-token prediction objective, predicting several future tokens per step rather than just the next one, and FP8 mixed-precision training, one of the first successful demonstrations of FP8 at this scale. DeepSeek pretrained V3 on 14.8 trillion tokens with a 128K context window, and disclosed the run's compute cost directly: about 2.79 million H800 GPU-hours. The widely repeated "$5.5-6 million" price tag is an analyst estimate applied to that GPU-hour figure, not a dollar amount DeepSeek itself published, worth keeping straight since the estimate is what triggered the market reaction, not an official number.

R1's training is arguably the more novel part. R1-Zero, an intermediate model, was trained with pure large-scale reinforcement learning directly on the V3 base checkpoint, no supervised fine-tuning warm-start at all, rewarding only final-answer correctness and output format. That process produced genuinely emergent behavior, the model learned to pause and re-check its own reasoning without being explicitly trained to do so, but also had readability problems and mixed languages within a single response. R1 itself fixes that by adding a small supervised "cold-start" dataset before RL, plus further SFT and RL stages on top. The RL algorithm behind both, Group Relative Policy Optimization (GRPO), originated in DeepSeek's earlier DeepSeekMath work: it drops the separate critic model that standard PPO needs and instead computes each output's advantage relative to a group of other sampled outputs for the same prompt, which is a meaningful part of why DeepSeek could run this at scale without a second model's worth of extra memory and compute. DeepSeek then distilled roughly 800,000 reasoning traces generated by R1 into six smaller dense models, Qwen-based at 1.5B/7B/14B/32B and Llama-based at 8B/70B, using supervised fine-tuning alone with no RL, released as the R1-Distill line.

Later releases kept iterating on the same base: V3.1 (August 2025) unified thinking and non-thinking modes into one model instead of splitting them across separate endpoints, and V3.2 (September-December 2025) added DeepSeek Sparse Attention for cheaper long-context inference alongside further agentic tool-use RL. V4 (April 2026) ships as two sizes, V4-Pro at 1.6 trillion total parameters with 49B active, and V4-Flash at 284B total with 13B active, both with a 1-million-token context window and reportedly pretrained on over 32 trillion tokens. V4's deeper architectural changes are documented mainly through Hugging Face model cards and third-party technical writeups rather than a peer-reviewed report the way V3 and R1 were, so treat its exact internals as reported rather than fully verified.

TRAINING EFFICIENCY

FP8 training, auxiliary-loss-free load balancing, and multi-token prediction are the specific techniques behind DeepSeek's disclosed compute figures, not just a smaller model.

MATH & REASONING

R1's pure-RL training pipeline and its successors have posted gold-medal-level results on competitions like the IMO, CMO, and ICPC World Finals.

DISTILLED FOR SELF-HOSTING

The R1-Distill dense models, from 1.5B to 70B, put R1-style reasoning behavior on ordinary Qwen and Llama-sized checkpoints that fit on a single GPU.

When to use it

DeepSeek's core pitch is frontier-comparable performance, especially on math, coding, and multi-step reasoning, at a fraction of the training and serving cost most labs assumed was required. By 2026, it sits alongside Kimi, Qwen, and Zhipu's GLM as one of several leading Chinese open-weight families rather than a singular outlier, each with a slightly different edge: DeepSeek is generally cited as strong on agentic coding and graduate-level reasoning, Kimi on long-horizon multi-step agent tasks, GLM on cost-efficient day-to-day coding, and Qwen on breadth and adoption. Against closed Western frontier models, DeepSeek's own benchmarks position V3.2 and V4 as competitive on reasoning at a much lower API price, though those comparisons come from DeepSeek's own marketing and are worth verifying independently rather than taken at face value.

In practice, the size decision matters more than the model choice. The full V3 or V4 flagship models, at 671B to 1.6 trillion total parameters, need real multi-node infrastructure to self-host, not a single card. If you want R1-style reasoning behavior without that infrastructure, the R1-Distill dense models at 1.5B through 70B are the practical target, they trade some of R1's ceiling for running on hardware you likely already have.

DeepSeek R1's distillation pipeline A diagram showing DeepSeek-R1's reasoning traces distilled through supervised fine-tuning alone into six smaller dense models, Qwen-based at 1.5 billion, 7 billion, 14 billion, and 32 billion parameters, and Llama-based at 8 billion and 70 billion parameters, as the practical self-hosting option compared to R1 itself. R1 ~800K TRACES SFT ONLY QWEN 1.5B / 7B / 14B / 32B LLAMA 8B / 70B R1-DISTILL SINGLE-GPU READY
Reasoning behavior distilled into existing dense architectures, not a shrunk version of R1's own MoE design.

How to use it

DeepSeek's model weights, including V3, R1, the R1-Distill line, and V4, are released under a permissive MIT license, allowing commercial use, modification, and self-hosting without royalty, one of the most straightforward licenses of any frontier- adjacent model family. Weights are published on Hugging Face under the deepseek-ai org, with inference code and technical reports on GitHub. For serving, vLLM and SGLang both typically ship support close to release day, SGLang in particular is notable for AMD GPU compatibility in both FP8 and BF16, and TensorRT-LLM is commonly supported for NVIDIA-optimized deployment. The full MoE flagship models require a real multi-GPU cluster to self-host given their total parameter counts; the R1-Distill dense models are the models actually meant for single or few-GPU self-hosting.

Beyond self-hosting, DeepSeek runs its own API platform, widely noted for aggressive pricing relative to comparable closed models, though exact rates have changed more than once and are worth checking directly against DeepSeek's own pricing page rather than a cached figure before budgeting around it. Fine-tuning support is broad through standard Hugging Face transformers tooling plus common community frameworks like Unsloth and Axolotl, with the distilled dense models again being the practical target for teams fine-tuning on their own data rather than the full-scale MoE checkpoints.

Where this fits at Numerata

Whether you're working with a full DeepSeek flagship or one of the R1-Distill models, P95 is where you fine-tune it on your own data, on infrastructure you control, private cloud or fully air-gapped, and NinetyFive serves it afterward, so a model built around efficient training also gets efficient serving, rather than losing that advantage to a naive deployment.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog