Meta releases Llama 4 Scout and Maverick - MoE models that run on consumer hardware

Meta released Llama 4 Scout (109B total, 17B active) and Maverick (400B total, 17B active) as open-weight MoE models. Scout runs on a single 48GB GPU with quantization, making it practical for local AI workstations.

Meta releases Llama 4 Scout and Maverick - MoE models that run on consumer hardware

Meta released two new open-weight models on April 5: Llama 4 Scout and Llama 4 Maverick. Both use a Mixture of Experts (MoE) architecture, which means they have large total parameter counts but only activate a fraction for each request.

Scout has 109 billion total parameters with 17 billion active per token. At 4-bit quantization, it needs roughly 24GB of VRAM, which means it runs on a single RTX 3090 or RTX 4090. Maverick is larger at 400 billion total with the same 17 billion active, requiring approximately 80GB of VRAM for quantized inference. That puts Maverick in dual-GPU or cloud territory.

⚡ New to this?

MoE (Mixture of Experts) means the model has many parameters total but only activates a small fraction for each request. This is like having a team of specialists where only the relevant expert answers each question. The result: big-model quality on smaller hardware.

🦞 OpenClaw angle

If you are building a local AI workstation with an RTX 3090 (24GB VRAM), Llama 4 Scout at 4-bit quantization fits in roughly 24GB. Download it via Ollama (ollama pull llama4-scout) and test it as a local agent backend. For dual-GPU setups with NVLink, Maverick becomes an option too.

The MoE approach is what makes these models interesting for self-hosters. Traditional dense models use all their parameters for every token, which means a 109B dense model would need far more VRAM than most people have. MoE gives you the quality of a larger model with the compute cost of a smaller one. The router network decides which expert subnetwork to activate for each token, keeping inference fast and memory usage manageable.

Ollama, vLLM, and llama.cpp all had day-one support. Quantized GGUF packs appeared on Hugging Face within hours of release, and Meta's own llama-stack provided an official deployment path immediately. Unsloth also shipped Llama 4 fine-tuning support shortly after, making customization practical on consumer hardware.

Over 1.2 million downloads in the first week made Scout one of the fastest-adopted open models ever. The community response was strongly positive, with r/LocalLLaMA discussions highlighting the quality-to-VRAM ratio as the best available for any open model.

The Llama 4 Community license allows commercial use, though it has different terms than Apache 2.0 or MIT. Revenue thresholds and usage conditions apply, so review the specifics if you plan to build a product on top of it. For personal projects and internal tools, the license is effectively unrestricted.

For anyone building local AI infrastructure, Scout is the model to test first. It hits a sweet spot between quality and hardware requirements that no other model matched in April. Maverick is the stretch goal for people with multi-GPU setups or cloud instances with high-VRAM configurations.

Source: Meta ↗

More from AI News