NVIDIA Nemotron 3 Ultra Comes to Kilo Code Free

NVIDIA’s Nemotron 3 Ultra is now available in Kilo Code, with free access for a limited time. The open-weights model was introduced by Jensen Huang at Computex 2026 and is being positioned for agentic coding, long-context work, and self-hosted deployment.

NVIDIA Nemotron 3 Ultra Comes to Kilo Code Free

NVIDIA’s Nemotron 3 Ultra is now available in Kilo Code, and Kilo says the model is free to use across its products for a limited time. The company announced the integration after NVIDIA CEO Jensen Huang introduced the model during his keynote at Computex 2026 in Taipei.

Kilo describes Nemotron 3 Ultra as NVIDIA’s flagship open-weights model and says it is aimed at agentic coding workflows. That means it is built for software agents that can plan tasks, call tools, navigate codebases, and generate structured code over multiple steps.

⚡ New to this?

This matters because an “open-weights” model is a large AI system where the model weights are available for use and customization, unlike fully closed products. For non-specialists, the big story is that NVIDIA is pushing a model aimed at coding agents that can handle longer tasks and larger projects.

A million-token context window means the model can keep a lot more code and documentation in memory during one session. That can help reduce context loss when working across big repos or long error trails.

🦞 OpenClaw angle

If you build self-hosted AI tools, test Nemotron 3 Ultra against your own codebases rather than relying on generic benchmarks. Focus on long-horizon workflows such as multi-file refactors, tool calls, and log analysis, because that is where Kilo says the model is meant to perform.

If you need data control, plan for an on-prem or isolated deployment path and review NVIDIA NeMo for customization and evaluation. If you already support Kilo Code users, add this model to your model-routing layer so you can compare its speed and planning quality against your current open model stack while the free access window is active.

According to Kilo, Nemotron 3 Ultra is a 550 billion-parameter model built on a hybrid Mamba-Transformer Mixture-of-Experts, or MoE, architecture. Kilo says only 55 billion parameters are active per token during inference, which is part of how the model is able to produce more than 300 tokens per second while still handling large reasoning tasks.

On stage at Computex, Huang said the model achieved a high PinchBench score and called it “frontier smart,” according to Kilo’s post. Kilo also says Nemotron 3 Ultra currently holds the “Best Open-Weights” title on the PinchBench agentic leaderboard with a 90% median success rate.

Kilo points to several reasons it expects the model to work well inside Kilo Code. The company says the model scored 48 on the Artificial Analysis Intelligence Index, which it describes as placing Nemotron 3 Ultra at the top of open models from the US in that ranking.

Another major feature is the context window. Kilo says Nemotron 3 Ultra natively supports up to 1,000,000 tokens of context, which would let users load entire codebases, long API references, and large error logs into a single session without running into short-context limits as quickly.

The company also says the model is designed for complex coding environments and was trained with contributions from the Nemotron Coalition. Kilo says the family has been optimized for multi-environment reinforcement learning, including SWE-RL, a training approach used to improve software engineering tasks through reinforcement learning.

Kilo contrasted Nemotron 3 Ultra with Nemotron 3 Super, a 120B-parameter open hybrid MoE model it released earlier this year. According to Kilo, Super became a daily driver for many users, but it had limits around planning and long-horizon tasks. The company says Ultra is intended to push further on those tasks.

Kilo also said Nemotron 3 Ultra performed similarly to Qwen 3.7 Plus in its internal KiloBench evaluations. The company said it will compare the models on the Kilo Leaderboard for agentic tasks such as coding and planning.

Beyond performance claims, Kilo is emphasizing deployment flexibility. The company says Nemotron models are released with open weights, datasets, and recipes, which gives organizations more transparency and control. Kilo says developers can use NVIDIA NeMo to customize, evaluate, and optimize the model, and that the family can be deployed in self-hosted environments for regulatory, sovereignty, or data-localization needs.

Kilo says Nemotron 3 Ultra is available now inside Kilo Code in terminal, VS Code, and JetBrains IDEs. The company also says the free offer is temporary.

Source: Kilo Blog ↗

More from AI News