Nemotron 3.5 Lightning
NVIDIA· released 11 Aug 2026
nvidia/nemotron-3.5-lightning- Input / 1M tokens
- $0.06
- Output / 1M tokens
- $0.16
Longest answer: 33K tokens
About
Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3 billion active parameters out of 30 billion total. It is aimed at high-throughput agentic workloads and specialized tasks. The model accepts text and returns text, with tools and reasoning listed as supported. Its context window is 262,144 tokens and its maximum output is 32,768 tokens. Open weights are available, with the Hugging Face ID nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16. The model ID is nvidia/nemotron-3.5-lightning. Listed prices are USD 0.06 per 1 million input tokens and USD 0.16 per 1 million output tokens. Teams considering it can assess its agentic focus alongside tool support, its text-only modality, context and output limits, and the stated per-token rates.
Who it is for
It suits teams building text-based, high-throughput agentic workloads or specialized tasks. Tool and reasoning support may fit workflows that rely on those capabilities.
What is good
- Open weights are available.
- Supports tools and reasoning.
- 262,144-token context window.
- Input costs USD 0.06 per 1 million tokens.
What to know first
- Inputs and outputs are text only.
- Maximum output is 32,768 tokens.
Verdict
Nemotron 3.5 Lightning combines open weights, tools, and reasoning with a focus on high-throughput agentic work. Its text-only modality and stated output limit are relevant boundaries when assessing fit.
Details
- Lab
- NVIDIAopenrouter.ai · 3 Oct 2026
- Context
- 262Kopenrouter.ai · 3 Oct 2026
- Input price
- $0.06 / 1Mopenrouter.ai · 3 Oct 2026
- Output price
- $0.16 / 1Mopenrouter.ai · 3 Oct 2026
- Max output
- 32,768 tokensopenrouter.ai · 3 Oct 2026
- Inputs
- textopenrouter.ai · 3 Oct 2026
- Open weights
- Yesopenrouter.ai · 3 Oct 2026
More newest models
See the listListed on Inferse
- Newest Models in 2026466 listed
- Longest-Context Models in 2026466 listed
- Cheapest Language Models in 2026437 listed
- Models with Tool Calling in 2026395 listed
- Reasoning Models in 2026332 listed
- Open-Weight Models in 2026178 listed
- NVIDIA Models: Pricing per 1M Tokens and Context (2026)11 listed
Sources
- openrouter.ai/nvidia/nemotron-3.5-lightning· checked 3 Oct 2026


