Nemotron 3 Ultra
NVIDIA· released 4 Jun 2026
nvidia/nemotron-3-ultra-550b-a55b- Input / 1M tokens
- $0.5
- Output / 1M tokens
- $2.2
Longest answer: 16K tokens
About
Nemotron 3 Ultra is NVIDIA’s open-weight reasoning and orchestration model, with 55 billion active parameters in a 550-billion-parameter mixture-of-experts design. Its hybrid architecture combines Transformer and Mamba components. It accepts text and produces text, with reasoning and tool support. The model has a 262,144-token context window and a maximum output of 16,384 tokens. Open weights are available under the Hugging Face identifier nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16. For API use, input costs $0.5 per 1M tokens and output costs $2.2 per 1M tokens. The listed release date is June 4, 2026. This configuration is identified as nvidia/nemotron-3-ultra-550b-a55b, and its official page is on OpenRouter. Its text-only input and output make the interface straightforward to account for when considering it for a stack, while the large total parameter count and MoE structure define its model profile.
Who it is for
Teams looking for an open-weight model for text reasoning, orchestration, or tool-supported workflows may consider it. Its 262,144-token context window and per-million-token input and output prices are relevant when estimating fit.
What is good
- Open weights with a listed Hugging Face identifier
- Reasoning and tool support
- 262,144-token context window
- Input and output prices are listed
What to know first
- Maximum output is 16,384 tokens
- Listed inputs and outputs are text only
- 55B active within 550B total parameters
Verdict
Nemotron 3 Ultra combines open weights, reasoning, and tool support in a text-only model. Its context window is 262,144 tokens, with input at $0.5 and output at $2.2 per 1M tokens.
Details
- Lab
- NVIDIAopenrouter.ai · 3 Oct 2026
- Context
- 262Kopenrouter.ai · 3 Oct 2026
- Input price
- $0.5 / 1Mopenrouter.ai · 3 Oct 2026
- Output price
- $2.2 / 1Mopenrouter.ai · 3 Oct 2026
- Max output
- 16,384 tokensopenrouter.ai · 3 Oct 2026
- Inputs
- textopenrouter.ai · 3 Oct 2026
- Open weights
- Yesopenrouter.ai · 3 Oct 2026
More newest models
See the listListed on Inferse
- Newest Models in 2026466 listed
- Longest-Context Models in 2026466 listed
- Cheapest Language Models in 2026437 listed
- Models with Tool Calling in 2026395 listed
- Reasoning Models in 2026332 listed
- Open-Weight Models in 2026178 listed
- NVIDIA Models: Pricing per 1M Tokens and Context (2026)11 listed
Sources
- openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b· checked 3 Oct 2026


