BLOG · SLM SERIES

Qwen Explained

Qwen (通义千问, "Tongyi Qianwen") is the open-weight model family built by Tongyi Lab, Alibaba Cloud's AI research group. Since its closed beta in April 2023, it has grown from a single 7B model into a coordinated family spanning roughly half a billion to hundreds of billions of parameters, released almost entirely under the Apache 2.0 license. It's one of a handful of open model families, alongside Llama, Mistral, and Gemma, that most teams building on open weights will run into. Here's who builds it, how it's actually trained, and which size fits which job.

Who builds Qwen

Qwen comes out of Tongyi Lab, the AI research arm inside Alibaba Cloud. Alibaba Cloud commercializes the family through its DashScope API and Model Studio (Bailian) platform, while the open-weight checkpoints ship straight to Hugging Face and ModelScope, Alibaba's own model hub, for anyone to download and run themselves. Alibaba has reported more than 90,000 enterprise adoptions of Qwen models within the first year of its open release program, and by 2026 commentary routinely describes it as one of Alibaba's most valuable AI assets, alongside its cloud and chip businesses. Releases have been frequent and iterative: Qwen (2023), Qwen1.5, Qwen2, Qwen2.5, Qwen2.5-Coder, Qwen3 (April 2025), and Qwen3.5 (February 2026), each widening language coverage, context length, and coding/math performance over the last.

Qwen release timeline A timeline showing successive Qwen releases from 2023 through early 2026, each widening language coverage, context length, and coding or math performance over the last. QWEN 2023 QWEN1.5 EARLY 2024 QWEN2 / 2.5 MID-LATE 2024 QWEN3 APRIL 2025 QWEN3.5 FEBRUARY 2026
Roughly one major generation per year, each widening languages, context, and coding/math strength.

How it's trained

Qwen3, the family's most recent generation with a full public technical report, ships as two architecture families sharing one training recipe. The dense line runs from 0.6B up to 32B parameters, all using grouped-query attention. The mixture-of-experts line adds a 30B-parameter model that only activates 3B parameters per token, and a 235B-parameter flagship that activates 22B, routing each token through 8 of 128 experts. Pretraining data has scaled with each generation: Qwen2 trained on roughly 7 trillion tokens, Qwen2.5 on 18 trillion across 29-plus languages, and Qwen3 on roughly 36 trillion tokens covering 119 languages and dialects.

The other notable design choice in Qwen3 is a single model that switches between a "thinking" mode, which reasons step by step before answering, and a fast "non-thinking" mode for simple queries, selectable at inference time instead of requiring two separate checkpoints. Post-training combines supervised fine-tuning with reasoning-focused reinforcement learning, and Alibaba distills the resulting reasoning ability from its larger checkpoints down into the smaller dense models, which is part of why the small end of the family punches above its parameter count on math and code benchmarks.

Qwen3.5, released in February 2026, moves the architecture on from there. Its open-weight flagship, Qwen3.5-397B-A17B, is built on what Alibaba calls the Qwen3-Next architecture: a higher-sparsity mixture-of-experts combined with a hybrid attention scheme (Gated DeltaNet plus gated attention) and multi-token prediction, which is what lets a 397-billion-parameter model activate only 17 billion parameters per token while decoding several times faster than the previous Qwen3-235B-A22B flagship at long context lengths. It's also natively multimodal, trained with early text-vision fusion rather than a vision tower bolted onto a text model afterward, and Alibaba widened language and dialect coverage from Qwen3's 119 to 201, alongside a larger 250K-token vocabulary that speeds up encoding and decoding across most of them. Post-training gains over Qwen3 came mostly from scaling up reinforcement learning across a much larger and harder set of tasks and environments, rather than tuning for specific benchmark metrics.

MULTILINGUAL

Qwen3 covered 119 languages and dialects; Qwen3.5 widened that to 201, a wider spread than most open model families, making it a common default for non-English or multilingual deployments.

CODE & MATH

A dedicated Qwen2.5-Coder line and reasoning-focused post-training give the family a strong reputation on coding and math benchmarks, at nearly every size tier.

AGENTIC & MULTIMODAL

Native tool-calling templates and the companion Qwen-Agent framework support building agents; Qwen3.5 adds native vision-language understanding on top of that, not a bolted-on vision module.

When to use it

Qwen's size range maps fairly directly onto deployment constraints. The 0.6B-4B dense models fit edge and on-device use, where memory is the binding constraint. The 8B-32B dense models are the common range for single-GPU production serving, and are where most fine-tuning work in the family happens. The 30B-A3B mixture-of-experts model is worth a look when throughput matters more than raw parameter count, since only 3B parameters activate per token despite the 30B total. For frontier-adjacent quality, or a native multimodal task, the current open-weight flagship is Qwen3.5-397B-A17B, which activates 17B of its 397B total parameters and needs real serving infrastructure to run well. If you'd rather not self-host something that large, Qwen3.5-Plus is the hosted equivalent through Alibaba Cloud Model Studio, with a 1M token context window and built-in tool use by default. Alibaba's "Max" tiers, similarly, are kept proprietary and API-only rather than released as open weights.

Reach for Qwen specifically when multilingual coverage, code generation, or math/reasoning are central to the task, that's where its training emphasis shows up most. For a general-purpose English chatbot or a task with a mature existing fine-tuning ecosystem around it, Llama remains a reasonable default; for the smallest viable footprint, Gemma's smaller checkpoints are worth comparing against Qwen's 0.6B-4B tier before committing.

Qwen size tiers mapped to deployment fit A chart mapping Qwen model sizes to deployment scenarios: 0.6B to 4B dense models for edge and on-device use, 8B to 32B dense models for single-GPU production serving, 30B-A3B mixture-of-experts for high throughput at low active-parameter cost, and the Qwen3.5-397B-A17B open-weight flagship for frontier-adjacent, natively multimodal quality. 0.6B-4B EDGE / ON-DEVICE 8B-32B SINGLE-GPU SERVING 30B-A3B HIGH THROUGHPUT 397B-A17B (QWEN3.5) MULTIMODAL FLAGSHIP
Size tier tracks deployment fit, from a single edge device up to a served MoE cluster.

How to use it

The open-weight Qwen3 lineup, dense and mixture-of-experts alike, ships under Apache 2.0, permissive enough for unrestricted commercial use. Weights are published on Hugging Face and ModelScope, with GGUF, AWQ, and AutoGPTQ quantized variants available from the community for constrained hardware. For serving, vLLM and SGLang are the standard choices for production throughput and support Qwen3's long context and thinking-mode switching; Ollama, LM Studio, and llama.cpp cover local and on-device use. Fine-tuning support is broad, standard Hugging Face transformers tooling and LoRA/QLoRA frameworks all work against the chat template out of the box, and the Qwen-Agent framework handles the tool-calling scaffolding if you're building an agent rather than a chat interface. If you'd rather try the models before self-hosting, Qwen Chat offers auto, thinking, and fast response modes, and the hosted Qwen3.5-Plus is reachable through Alibaba Cloud Model Studio's DashScope API with an enable_thinking flag for chain-of-thought reasoning and an enable_search flag for built-in web search and code execution.

The main thing to check per-release rather than assume: Alibaba's largest "Max" tier is typically proprietary and API-only at launch, with an open-weight version sometimes following later, so verify what's actually downloadable for the specific release and size you're targeting before planning a deployment around it.

Where this fits at Numerata

Whichever Qwen size you land on, the same before/after split applies as with any open-weight model: P95 is where you fine-tune it on your own data, on infrastructure you control, private cloud or fully air-gapped, and NinetyFive is what serves it afterward, turning that size tier's parameter count into the sub-50ms latency the architecture is actually capable of.

Numerata runs inside your own environment: private cloud, on-prem, or fully air-gapped.  ·  Back to blog