Meta releases Llama 4 Scout and Maverick with 17B active parameters

Meta shipped two Mixture-of-Experts models in the Llama 4 family. Scout has 109B total parameters with 17B active and fits on a single 48GB GPU at 4-bit quantization. Maverick scales to 400B total with the same 17B active count for heavier workloads. Both available on Hugging Face and Ollama with day-one community tooling from Unsloth and vLLM.

Meta releases Llama 4 Scout and Maverick with 17B active parameters

Meta has released two new models in the Llama 4 family, Scout and Maverick, and both use a Mixture-of-Experts design rather than a single dense network. According to Meta’s announcement, Scout has 109 billion total parameters with 17 billion active at inference time, while Maverick scales up to 400 billion total parameters with the same 17 billion active count.

That active-parameter number is the part that matters most for deployment. In a Mixture-of-Experts model, the system contains multiple sub-networks, but only a subset is used for any given token or request. That lets model builders increase overall capacity without forcing every query through the full parameter count, which is one reason these models are getting attention from people running AI systems on local hardware.

⚡ New to this?

This is a new release of large language models, the kind of AI systems that generate text, answer questions, and power assistants or agents. A Mixture-of-Experts model is a way to make a model bigger without using all of its parts at once, which can make it cheaper and easier to run. People who build AI tools care because these models may be easier to self-host and integrate into software than older, fully dense models.

🦞 OpenClaw angle

If you self-host models via Ollama or vLLM, pull Llama 4 Scout and test it as an agent backbone. At 17B active parameters it is fast enough for tool-calling workflows while being far more capable than smaller dense models. Quantized GGUF files were available on Hugging Face within hours of launch.

Meta says Scout is designed to fit on a single 48GB GPU when run at 4-bit quantization. Quantization reduces the precision of model weights so the model uses less memory, which can make large models practical on smaller systems, at some cost to quality depending on the setup. For teams that self-host, that puts a model in the Llama 4 family into a range that is easier to test without a large multi-GPU cluster.

Maverick uses the same 17 billion active parameter count but pushes the total parameter budget much higher, making it the larger option for heavier workloads. Meta has not presented this as a small-model release; the point is to give developers a choice between a more compact model that is easier to run and a larger sibling aimed at more demanding use cases.

The company made both models available through Hugging Face and Ollama on launch day, which matters because those ecosystems are where many developers first try new open-weight models. Hugging Face is the main hub for model files and community distribution, while Ollama is a common local-running tool for people who want to test models on their own machines.

Meta also had day-one support from familiar tooling in the open model ecosystem. Unsloth, which is used for fine-tuning and training efficiency, and vLLM, a popular inference server for high-throughput model serving, were both part of the initial launch story. That kind of early tooling support often determines how quickly a model moves from announcement to real use.

The release continues Meta’s push to keep Llama competitive with other frontier model families while also making it easier for independent developers and enterprise teams to work with the models directly. For the open-weight community, the combination of a 17B active-parameter design, broad distribution, and immediate tooling support is usually more important than raw size alone.

Scout is the model Meta is positioning for the most practical local and cloud deployments, especially where memory headroom is limited. Maverick, with its larger total parameter count, gives a clearer path for higher-capacity workloads without changing the active compute footprint that much at inference time.

Meta said both models are part of the Llama 4 family, and the release landed with code and model distribution already set up across the most common places developers look first, including Hugging Face, Ollama, Unsloth, and vLLM.

Source: Meta AI ↗

More from AI News