Nemotron 3.5 Lightning

NVIDIA· released 11 Aug 2026

nvidia/nemotron-3.5-lightning
Input / 1M tokens
$0.06
Output / 1M tokens
$0.16
Context window262K tokens

Longest answer: 33K tokens

reads textopen weightstool callingreasoning

About

Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3 billion active parameters out of 30 billion total. It is aimed at high-throughput agentic workloads and specialized tasks. The model accepts text and returns text, with tools and reasoning listed as supported. Its context window is 262,144 tokens and its maximum output is 32,768 tokens. Open weights are available, with the Hugging Face ID nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16. The model ID is nvidia/nemotron-3.5-lightning. Listed prices are USD 0.06 per 1 million input tokens and USD 0.16 per 1 million output tokens. Teams considering it can assess its agentic focus alongside tool support, its text-only modality, context and output limits, and the stated per-token rates.

Who it is for

It suits teams building text-based, high-throughput agentic workloads or specialized tasks. Tool and reasoning support may fit workflows that rely on those capabilities.

What is good

  • Open weights are available.
  • Supports tools and reasoning.
  • 262,144-token context window.
  • Input costs USD 0.06 per 1 million tokens.

What to know first

  • Inputs and outputs are text only.
  • Maximum output is 32,768 tokens.

Verdict

Nemotron 3.5 Lightning combines open weights, tools, and reasoning with a focus on high-throughput agentic work. Its text-only modality and stated output limit are relevant boundaries when assessing fit.

Details

Lab
NVIDIAopenrouter.ai · 3 Oct 2026
Context
262Kopenrouter.ai · 3 Oct 2026
Input price
$0.06 / 1Mopenrouter.ai · 3 Oct 2026
Output price
$0.16 / 1Mopenrouter.ai · 3 Oct 2026
Max output
32,768 tokensopenrouter.ai · 3 Oct 2026
Inputs
textopenrouter.ai · 3 Oct 2026
Open weights
Yesopenrouter.ai · 3 Oct 2026

More newest models

See the list