NVIDIA releases Nemotron 3 Nano Omni reasoning model

NVIDIA has released Nemotron 3 Nano Omni, a 31B multimodal model for video, audio, image and text tasks. The company says it is available for commercial use and can be run with vLLM, SGLang, TensorRT-LLM, llama.cpp and Ollama on supported NVIDIA GPUs.

NVIDIA releases Nemotron 3 Nano Omni reasoning model

NVIDIA has released Nemotron 3 Nano Omni, a multimodal model that can process video, audio, images and text, according to the model page published on 04/28/2026. The company says the model is designed for enterprise workloads such as summarization, transcription, document intelligence and GUI automation.

The model is listed as Nemotron-3-Nano-Omni-30B-A3B-Reasoning, with about 31 billion parameters and a Mamba2-Transformer hybrid mixture-of-experts architecture. NVIDIA says it combines a Nemotron 3 Nano LLM with a CRADIO v4-H vision encoder and a Parakeet speech encoder.

⚡ New to this?

This news matters because NVIDIA is shipping a single model that can read text, look at images, listen to audio and analyze video. For a non-expert, that means one AI system can handle things like meeting recordings, scanned documents and screen-based workflows instead of using separate tools for each format.

The model also includes reasoning and tool-calling support, which means it can think through a task and trigger other software actions. That is important for teams building AI assistants, document processors or automation systems.

🦞 OpenClaw angle

If you run self-hosted agents, treat this as a multimodal backend, not just a chat model. Start by testing it against one concrete workflow, such as document intake or browser-based GUI automation, and measure whether its JSON output and tool-calling stay consistent under your prompts.

If you use video or audio, set frame sampling and audio handling explicitly at serve time instead of relying on defaults. Also build a small validation set from your own data before deployment, since NVIDIA says use-case-specific testing is needed to verify safety and performance.

According to the documentation, the model supports up to 256k tokens of context, accepts video, audio, image and text inputs, and can output plain text, JSON, reasoning traces and tool calls. It also supports word-level timestamps for transcription. NVIDIA says the model is English-only.

NVIDIA lists several use cases for the model. These include customer service workflows such as verifying a delivery photo or drive-through order, media and entertainment tasks like dense captions and video search, document intelligence for contracts and financial documents, and GUI automation for AI agents that work in incident management, browser tasks and email.

The model page says Nemotron 3 Nano Omni was improved using several other models, including Qwen3-VL-30B-A3B-Instruct, Qwen3.5-122B-A10B, Qwen3.5-397B-A17B, Qwen2.5-VL-72B-Instruct and gpt-oss-120b. NVIDIA says the model is governed by the NVIDIA Open Model Agreement and is available for commercial use.

The release also includes deployment instructions. NVIDIA says the model works with vLLM, NeMo, Megatron, NeMo-RL, TensorRT-LLM, TensorRT Edge-LLM, llama.cpp, Ollama and SGLang, with support for NVIDIA Ampere, Hopper, Blackwell and Lovelace GPUs on Linux. The company provides separate setup notes for DGX Spark and Jetson Thor.

For vLLM, NVIDIA says version 0.20.0 is required. The documentation points users to BF16, FP8 and NVFP4 weight files on Hugging Face, and says the model file is about 62 GB, requiring at least 70 GB of free disk space for download.

NVIDIA also documents reasoning behavior and request settings. The model supports a thinking mode with a reasoning budget, and users can disable reasoning output by setting chat_template_kwargs to disable thinking. The page also recommends explicit video frame sampling settings, since the default can be too conservative for real video workloads.

For audio, NVIDIA says users need to install vLLM audio extras before serving the model. For PDF workflows, the company says the API does not accept raw PDF files and pages must be rendered to images first.

The training section says Nemotron-Omni was built on 354,587,705 data points, or about 717 billion tokens, across 1,395 dataset entries. NVIDIA says the data covers text, audio, image, video and mixed multimodal combinations, with additional curated examples for document reasoning, computer use and long-horizon workflows.

NVIDIA says the model can be embedded as an API call into the software stack described in the documentation. The company also warns that foundation and fine-tuned models should be tested with use-case-specific data before deployment, using iterative validation at unit and system level.

The new model is available globally, and NVIDIA says it can be used commercially under its open model agreement.

Source: r/LocalLLaMA ↗

More from AI News