Nvidia releases Nemotron 3 Nano Omni for edge AI agents
Nvidia has released Nemotron 3 Nano Omni, an open-weight multimodal model that handles vision, audio, and language in one architecture. The company says it runs on a single GPU, tops six benchmarks, and is available for commercial use under Nvidia’s Open Model Agreement.
Nvidia on Tuesday released Nemotron 3 Nano Omni, an open-weight multimodal AI model built to run autonomous AI agents on edge devices.
The model is designed to handle vision, audio, and language inside a single architecture. According to Nvidia, that lets one model take on tasks that often require separate systems for speech, image, video, and document processing.
Nvidia said Nemotron 3 Nano Omni has 30 billion parameters, but only 3 billion are active in each inference step. That design uses a mixture-of-experts approach, which routes each token to a small set of specialists rather than activating the full model every time.
The company said that makes the model efficient enough to run on a single GPU. Nvidia also said the model delivers nine times higher throughput than comparable open multimodal models with similar interactivity, 2.9 times faster single-stream reasoning on multimodal tasks, and about nine times greater effective system capacity for video reasoning.
Nvidia said the model can process text, images, audio, video, documents, charts, and graphical interfaces as input, and generate text as output. In practice, that means it is aimed at systems that need to see, hear, and read in one workflow, such as voice agents, document tools, video understanding systems, and computer-use agents.
The company said Nemotron 3 Nano Omni tops six benchmarks covering document intelligence, video understanding, and audio comprehension. Nvidia did not list the benchmarks in the source material, but it positioned the model as a strong performer across those categories.
Under the hood, Nvidia said the model uses a hybrid Mamba-Transformer architecture. The system includes 23 Mamba-2 selective state-space layers, 23 mixture-of-experts layers with 128 experts, and six grouped-query attention layers.
For images, the model uses the C-RADIOv4-H vision encoder, which handles variable-resolution inputs. For audio, Nvidia said it uses the Parakeet-TDT-0.6B-v2 encoder for speech and environmental sound. For video, the model uses three-dimensional convolutions to capture motion across frames rather than treating each frame as a separate image.
Nvidia said the base text model was pretrained on 25 trillion tokens and supports a 256,000-token context window. That gives it room to handle long inputs, including large documents and extended multimodal sessions.
The company said the architecture is meant to maximize capability per active parameter, since edge deployments are limited by inference compute rather than model size at rest. Nvidia also said the model can run on hardware it announced at GTC 2026, including DGX Spark and DGX Station workstations, without requiring the larger multi-GPU clusters used by bigger models in data centers.
Nemotron 3 Nano Omni is available on Hugging Face under Nvidia’s Open Model Agreement, which allows commercial use. Nvidia also said the model is available as a NIM microservice, through Amazon SageMaker JumpStart, and on OpenRouter.
The company listed deployment support for vLLM, SGLang, Ollama, llama.cpp, and TensorRT-LLM. Nvidia said that breadth is meant to make the model easier to adopt across different stacks and environments.
Nvidia said enterprise users already include Foxconn, Palantir, Aible, ASI, Eka Care, and H Company. It added that Dell, DocuSign, Infosys, Oracle, and Zefr are evaluating the model for production use.
The release marks a bigger shift for Nvidia, which has long sold the GPUs, networking, and software used to run AI systems. With Nemotron 3 Nano Omni, the company is also pushing into the model layer itself, and into the market for AI agents that need local, multimodal inference on Nvidia hardware.