Qwen 3 family ships eight model sizes - 72B beats GPT-4o on reasoning benchmarks
Alibaba's Qwen team released the Qwen 3 family with eight model sizes from 0.6B to 72B parameters. The 72B model became the first open-weight model to beat GPT-4o on MMLU-Pro, while Qwen3-Coder-32B offers 128K context with native tool calling.
Alibaba's Qwen team released the Qwen 3 model family in April, offering eight sizes ranging from 0.6B to 72B parameters. The lineup covers everything from edge devices to high-end inference servers, giving builders a model for nearly every hardware constraint.
The headline result is that Qwen3-72B became the first open-weight model to beat GPT-4o on MMLU-Pro, a challenging multi-task benchmark that tests graduate-level reasoning across multiple domains. This is significant because it shows open-weight models can now match proprietary offerings on standardized tests, not just in cherry-picked examples or narrow benchmarks.
Qwen3-Coder-32B is the standout for builders. It has a 128K context window and native tool calling, which means it can read large codebases and invoke functions without additional prompting tricks. At 32B parameters, it fits on a single GPU with quantization, making it practical for local development. The tool-calling capability is particularly important for agent frameworks that rely on function calling for MCP integration.
The eight sizes in the family are: 0.6B, 1.7B, 4B, 8B, 14B, 32B, 72B dense, and the Coder-32B variant. The 0.6B runs on devices with 2GB of RAM. The 4B and 8B models work well on laptops. The 14B and 32B models fit on single consumer GPUs. The 72B needs a multi-GPU setup or cloud instance with substantial VRAM.
All models are available on Hugging Face and Ollama with day-one support. The Qwen team provided quantized versions in multiple formats, and the community contributed additional GGUF packs within hours. The licensing allows commercial use, though the specific terms differ from Western open-source licenses, so review them before building commercial products.
For agent operators, the Coder variant deserves immediate testing. Native tool calling means the model can work with MCP servers and function-calling APIs without the wrapper hacks that smaller models often need. The 128K context window is large enough to hold an entire project's source code in a single request.
The Qwen 3 release reinforces a pattern from April 2026: open-weight models are arriving within weeks of frontier releases, at competitive quality levels, and runnable on hardware most builders already own.