ChengRang

NVIDIA Nemotron 3.5 Lightning

AI Platforms Free

NVIDIA's open-weight hybrid model with 30B total and roughly 3B active parameters, up to 1M-token context and tool calling, built for high-frequency agent execution steps and local deployment

NVIDIAMoEMambaOpen SourceTool Calling
Visit NVIDIA Nemotron 3.5 Lightning

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

Nemotron 3.5 Lightning is an open-weight hybrid model from NVIDIA with 30 billion total parameters and roughly 3 billion activated per token. It chains together Mamba-2, attention and mixture-of-experts layers, was pretrained on more than 20 trillion tokens, and pairs an NVFP4 recipe with multi-token prediction for faster generation. Context reaches up to 1M tokens and the model supports both reasoning and tool calling, covering English, coding languages, plus Spanish, French, German, Italian and Japanese. Weights are released under the OpenMDW License Agreement v1.1.

Its positioning is not standalone. NVIDIA pairs it with Nemotron 3 Ultra: Ultra is the heavy model at 550 billion total and 55 billion active parameters for tasks needing long reasoning chains, while Lightning takes on tool calling, coding and structured output, the high-frequency agent execution steps with clear boundaries, and works well as a resident execution unit for sub-agents.

The BF16 benchmarks NVIDIA publishes are MMLU Pro 81.94, GPQA Diamond 75.44, SWE-bench Verified 51.56, IFBench 71.88, AA-LCR 52.00 and Terminal-Bench 2.1 at 24.58, all on the vendor side of the ledger. Hosted API list prices sit around 0.07 to 0.08 dollars per million input tokens and 0.20 dollars for output, and OpenRouter published an explainer on September 22 covering its role in high-frequency agent execution calls.

Local deployment is the other selling point. A third-party test measured 334 tokens per second on an RTX 5090 with a first token after a 128k document at 14.9 seconds; on an RTX 3090 running a quantized GGUF build it reached 193.7 tokens per second, and moving context from 32k to 192k added only 1,119 MiB of VRAM. That memory efficiency at long context is what lets it run on consumer cards.

Key Features

Use Cases

Pros

Pricing

Hosted API list prices run about 0.07 to 0.08 dollars per million input tokens and 0.20 dollars for output, with some gateways offering it on a free tier. Weights are open under the OpenMDW License Agreement v1.1 for self-hosting. Check individual providers for current rates.

Summary

Nemotron 3.5 Lightning is an open-weight hybrid model from NVIDIA at 30 billion total and roughly 3 billion active parameters, combining Mamba-2, attention and mixture-of-experts layers, pretrained on more than 20 trillion tokens, with up to a 1M-token context, reasoning and tool calling, and weights under the OpenMDW license. NVIDIA pairs it with the 550B Nemotron 3 Ultra so it can take on high-frequency execution steps like tool calling and coding. Hosted API pricing sits around 0.07 to 0.08 dollars per million input tokens, and third-party consumer-card tests measured several hundred tokens per second with only a small memory increase at long context.

Version History

Category
AI Platforms
Pricing
Free
Tags
NVIDIA · MoE · Mamba
Website

Related Tools