Qwen 3 ships in eight sizes from 0.6B to 72B, beats GPT-4o on key benchmarks

Alibaba's Qwen team released the Qwen 3 model family in eight sizes from 0.6B to 72B dense plus a 235B MoE variant. The 72B flagship became the first open-weight model to beat GPT-4o on MMLU-Pro. The code-specialized Qwen3-Coder-32B offers 128K context and native tool calling, pulling 420,000 downloads in its first week.

Qwen 3 ships in eight sizes from 0.6B to 72B, beats GPT-4o on key benchmarks

Alibaba’s Qwen team has released Qwen 3, a new open-weight model family that spans eight dense sizes from 0.6 billion parameters to 72 billion, plus a larger 235 billion parameter mixture-of-experts variant. The launch puts a broad range of model sizes into one lineup, from small systems that can run on modest hardware to large models aimed at higher-end inference and benchmark performance.

The headline result is the 72B flagship model, which the Qwen team says became the first open-weight model to beat GPT-4o on MMLU-Pro. MMLU-Pro is a harder version of the widely used MMLU benchmark, designed to test reasoning and knowledge across a broad mix of subjects. Open-weight means the model’s weights are available to download and run, which is different from a closed model that can only be accessed through a vendor’s hosted API.

⚡ New to this?

Qwen 3 is a new family of AI models from Alibaba’s Qwen team. Open-weight means the model files can be downloaded and run outside the company’s own service, which matters to teams that self-host AI or need more control over where data goes.

The benchmark claim is notable because MMLU-Pro is a common test for broad knowledge and reasoning, and GPT-4o is one of the best-known closed models. Qwen3-Coder-32B also matters because a code model with a 128K context window can read much longer inputs at once, which is useful for software projects and automation tools.

🦞 OpenClaw angle

Test Qwen3-Coder-32B as an alternative agent model for coding and tool-calling tasks. Its 128K context window and native function calling support make it a strong candidate for self-hosted agent workflows. Available via Ollama with ollama pull qwen3:32b.

According to the Qwen blog, the family is built as a reasoning-oriented system with support for both concise responses and more deliberate step-by-step behavior. That fits a broader trend among frontier models, where vendors are trying to balance chat performance, coding ability, and agent-like task execution in the same base model line.

The size spread matters because it gives developers more room to match model capacity to deployment constraints. Smaller models can be easier to self-host, cheaper to run, and more practical for latency-sensitive applications, while the larger versions are aimed at tasks where benchmark performance and reasoning depth matter more.

Qwen also introduced Qwen3-Coder-32B, a code-specialized model in the same family. The company says it supports a 128K context window and native tool calling, two features that are especially relevant for software agents that need to work across long repositories, documentation, logs, or multi-step workflows.

A 128K context window lets the model process a much larger amount of text in one go than older systems with shorter context limits. Native tool calling means the model can be trained or configured to request external actions, such as calling a function, querying a database, or using another software tool, instead of only producing plain text.

The coder model also appears to have attracted early interest. Qwen said Qwen3-Coder-32B reached 420,000 downloads in its first week, which suggests strong demand for a self-hostable coding model that can handle longer inputs and structured tool use.

The broader Qwen 3 release continues Alibaba’s push to compete in both open and enterprise-friendly AI. The family now covers a wide range of deployment scenarios, from compact models intended for local or low-cost use to large-scale systems that can be tested against proprietary leaders on standard benchmarks.

For teams building agents, copilots, or internal automation systems, the important part is not just model size. It is the combination of open weights, multiple deployment tiers, longer context, and tool integration, all of which are increasingly central to how modern AI systems are built and deployed.

Source: Qwen (Alibaba) ↗

More from AI News