Nvidia releases Nemotron 3 Nano Omni for edge AI agents

Nvidia has released Nemotron 3 Nano Omni, an open-weight multimodal model that handles vision, audio, and language in one architecture. The company says it runs on a single GPU, tops six benchmarks, and is available for commercial use under Nvidia’s Open Model Agreement.

Nvidia releases Nemotron 3 Nano Omni for edge AI agents

Nvidia on Tuesday released Nemotron 3 Nano Omni, an open-weight multimodal AI model built to run autonomous AI agents on edge devices.

The model is designed to handle vision, audio, and language inside a single architecture. According to Nvidia, that lets one model take on tasks that often require separate systems for speech, image, video, and document processing.

⚡ New to this?

This matters because a multimodal model can work with several types of data at once, such as text, images, audio, and video. A GPU, or graphics processing unit, is the hardware these models run on, and Nvidia is saying this one can run on a single GPU instead of a large server cluster. For companies building AI systems, that can mean simpler setups and lower infrastructure demands.

🦞 OpenClaw angle

If you build self-hosted agents, test whether a single multimodal model can replace separate OCR, speech-to-text, and vision services in your stack. Start by benchmarking Nemotron 3 Nano Omni on one workflow that mixes text plus images or audio, then compare latency and memory use against your current pipeline.

If you already deploy on Nvidia hardware, try it through vLLM, Ollama, or TensorRT-LLM first so you can measure integration cost before rewriting anything. Keep your model routing abstracted behind one endpoint, because this release shows Nvidia is pushing a full-stack path where the model, runtime, and hardware are all tied together.

Nvidia said Nemotron 3 Nano Omni has 30 billion parameters, but only 3 billion are active in each inference step. That design uses a mixture-of-experts approach, which routes each token to a small set of specialists rather than activating the full model every time.

The company said that makes the model efficient enough to run on a single GPU. Nvidia also said the model delivers nine times higher throughput than comparable open multimodal models with similar interactivity, 2.9 times faster single-stream reasoning on multimodal tasks, and about nine times greater effective system capacity for video reasoning.

Nvidia said the model can process text, images, audio, video, documents, charts, and graphical interfaces as input, and generate text as output. In practice, that means it is aimed at systems that need to see, hear, and read in one workflow, such as voice agents, document tools, video understanding systems, and computer-use agents.

The company said Nemotron 3 Nano Omni tops six benchmarks covering document intelligence, video understanding, and audio comprehension. Nvidia did not list the benchmarks in the source material, but it positioned the model as a strong performer across those categories.

Under the hood, Nvidia said the model uses a hybrid Mamba-Transformer architecture. The system includes 23 Mamba-2 selective state-space layers, 23 mixture-of-experts layers with 128 experts, and six grouped-query attention layers.

For images, the model uses the C-RADIOv4-H vision encoder, which handles variable-resolution inputs. For audio, Nvidia said it uses the Parakeet-TDT-0.6B-v2 encoder for speech and environmental sound. For video, the model uses three-dimensional convolutions to capture motion across frames rather than treating each frame as a separate image.

Nvidia said the base text model was pretrained on 25 trillion tokens and supports a 256,000-token context window. That gives it room to handle long inputs, including large documents and extended multimodal sessions.

The company said the architecture is meant to maximize capability per active parameter, since edge deployments are limited by inference compute rather than model size at rest. Nvidia also said the model can run on hardware it announced at GTC 2026, including DGX Spark and DGX Station workstations, without requiring the larger multi-GPU clusters used by bigger models in data centers.

Nemotron 3 Nano Omni is available on Hugging Face under Nvidia’s Open Model Agreement, which allows commercial use. Nvidia also said the model is available as a NIM microservice, through Amazon SageMaker JumpStart, and on OpenRouter.

The company listed deployment support for vLLM, SGLang, Ollama, llama.cpp, and TensorRT-LLM. Nvidia said that breadth is meant to make the model easier to adopt across different stacks and environments.

Nvidia said enterprise users already include Foxconn, Palantir, Aible, ASI, Eka Care, and H Company. It added that Dell, DocuSign, Infosys, Oracle, and Zefr are evaluating the model for production use.

The release marks a bigger shift for Nvidia, which has long sold the GPUs, networking, and software used to run AI systems. With Nemotron 3 Nano Omni, the company is also pushing into the model layer itself, and into the market for AI agents that need local, multimodal inference on Nvidia hardware.

Source: The Next Web ↗

More from AI News