Qwen (通义千问, "Tongyi Qianwen") is the open-weight model family
built by Tongyi Lab, Alibaba Cloud's AI research group. Since its
closed beta in April 2023, it has grown from a single 7B model into
a coordinated family spanning roughly half a billion to hundreds of
billions of parameters, released almost entirely under the Apache
2.0 license. It's one of a handful of open model families, alongside
Llama, Mistral, and Gemma, that most teams building on open weights
will run into. Here's who builds it, how it's actually trained,
and which size fits which job.
Who builds Qwen
Qwen comes out of Tongyi Lab, the AI research
arm inside Alibaba Cloud. Alibaba Cloud commercializes the family
through its DashScope API and Model Studio (Bailian) platform, while
the open-weight checkpoints ship straight to Hugging Face and
ModelScope, Alibaba's own model hub, for anyone to download and run
themselves. Alibaba has reported more than 90,000 enterprise
adoptions of Qwen models within the first year of its open release
program, and by 2026 commentary routinely describes it as one of
Alibaba's most valuable AI assets, alongside its cloud and chip
businesses. Releases have been frequent and iterative: Qwen (2023),
Qwen1.5, Qwen2, Qwen2.5, Qwen2.5-Coder, Qwen3 (April 2025), and
Qwen3.5 (February 2026), each widening language coverage, context
length, and coding/math performance over the last.
Roughly one major generation per year, each widening languages, context, and coding/math strength.
How it's trained
Qwen3, the family's most recent generation with a full public
technical report, ships as two architecture families sharing one
training recipe. The dense line runs from 0.6B up
to 32B parameters, all using grouped-query attention. The
mixture-of-experts line adds a 30B-parameter model
that only activates 3B parameters per token, and a 235B-parameter
flagship that activates 22B, routing each token through 8 of 128
experts. Pretraining data has scaled with each generation: Qwen2
trained on roughly 7 trillion tokens, Qwen2.5 on 18 trillion across
29-plus languages, and Qwen3 on roughly 36 trillion tokens covering
119 languages and dialects.
The other notable design choice in Qwen3 is a single model that
switches between a "thinking" mode, which reasons step by step
before answering, and a fast "non-thinking" mode for simple queries,
selectable at inference time instead of requiring two separate
checkpoints. Post-training combines supervised fine-tuning with
reasoning-focused reinforcement learning, and Alibaba distills the
resulting reasoning ability from its larger checkpoints down into
the smaller dense models, which is part of why the small end of the
family punches above its parameter count on math and code
benchmarks.
Qwen3.5, released in February 2026, moves the architecture on
from there. Its open-weight flagship, Qwen3.5-397B-A17B,
is built on what Alibaba calls the Qwen3-Next architecture: a
higher-sparsity mixture-of-experts combined with a hybrid attention
scheme (Gated DeltaNet plus gated attention) and multi-token
prediction, which is what lets a 397-billion-parameter model
activate only 17 billion parameters per token while decoding
several times faster than the previous Qwen3-235B-A22B flagship at
long context lengths. It's also natively multimodal, trained with
early text-vision fusion rather than a vision tower bolted onto a
text model afterward, and Alibaba widened language and dialect
coverage from Qwen3's 119 to 201, alongside a larger 250K-token
vocabulary that speeds up encoding and decoding across most of
them. Post-training gains over Qwen3 came mostly from scaling up
reinforcement learning across a much larger and harder set of
tasks and environments, rather than tuning for specific benchmark
metrics.
MULTILINGUAL
Qwen3 covered 119 languages and dialects; Qwen3.5 widened that to 201, a wider spread than most open model families, making it a common default for non-English or multilingual deployments.
CODE & MATH
A dedicated Qwen2.5-Coder line and reasoning-focused post-training give the family a strong reputation on coding and math benchmarks, at nearly every size tier.
AGENTIC & MULTIMODAL
Native tool-calling templates and the companion Qwen-Agent framework support building agents; Qwen3.5 adds native vision-language understanding on top of that, not a bolted-on vision module.
When to use it
Qwen's size range maps fairly directly onto deployment
constraints. The 0.6B-4B dense models fit edge and on-device use,
where memory is the binding constraint. The 8B-32B dense models are
the common range for single-GPU production serving, and are where
most fine-tuning work in the family happens. The 30B-A3B
mixture-of-experts model is worth a look when throughput matters
more than raw parameter count, since only 3B parameters activate per
token despite the 30B total. For frontier-adjacent quality, or a
native multimodal task, the current open-weight flagship is
Qwen3.5-397B-A17B, which activates 17B of its 397B total
parameters and needs real serving infrastructure to run well. If
you'd rather not self-host something that large, Qwen3.5-Plus is
the hosted equivalent through Alibaba Cloud Model Studio, with a 1M
token context window and built-in tool use by default. Alibaba's
"Max" tiers, similarly, are kept proprietary and API-only rather
than released as open weights.
Reach for Qwen specifically when multilingual coverage, code
generation, or math/reasoning are central to the task, that's where
its training emphasis shows up most. For a general-purpose English
chatbot or a task with a mature existing fine-tuning ecosystem
around it, Llama remains a reasonable default; for the smallest
viable footprint, Gemma's smaller checkpoints are worth comparing
against Qwen's 0.6B-4B tier before committing.
Size tier tracks deployment fit, from a single edge device up to a served MoE cluster.
How to use it
The open-weight Qwen3 lineup, dense and mixture-of-experts alike,
ships under Apache 2.0, permissive enough for
unrestricted commercial use. Weights are published on Hugging Face
and ModelScope, with GGUF, AWQ, and AutoGPTQ quantized variants
available from the community for constrained hardware. For serving,
vLLM and SGLang are the standard choices for production throughput
and support Qwen3's long context and thinking-mode switching;
Ollama, LM Studio, and llama.cpp cover local and on-device use.
Fine-tuning support is broad, standard Hugging Face
transformers tooling and LoRA/QLoRA frameworks all
work against the chat template out of the box, and the
Qwen-Agent framework handles the tool-calling scaffolding if you're
building an agent rather than a chat interface. If you'd rather try
the models before self-hosting, Qwen Chat offers auto, thinking,
and fast response modes, and the hosted Qwen3.5-Plus is reachable
through Alibaba Cloud Model Studio's DashScope API with an
enable_thinking flag for chain-of-thought reasoning
and an enable_search flag for built-in web search and
code execution.
The main thing to check per-release rather than assume: Alibaba's
largest "Max" tier is typically proprietary and API-only at launch,
with an open-weight version sometimes following later, so verify
what's actually downloadable for the specific release and size
you're targeting before planning a deployment around it.
Where this fits at Numerata
Whichever Qwen size you land on, the same before/after split
applies as with any open-weight model: P95
is where you fine-tune it on your own data, on infrastructure you
control, private cloud or fully air-gapped, and
NinetyFive is what serves it afterward,
turning that size tier's parameter count into the sub-50ms latency
the architecture is actually capable of.