Nemotron (NVIDIA) Pricing: Which Model Is Worth Paying For?
OpenRouter lists 10 Nemotron (NVIDIA) models. For 1M input plus 1M output tokens they range from $0.23 (Nemotron 3.5 Lightning) to $2.70 (Nemotron 3 Ultra), a 12x spread. Most teams overpay by defaulting to the flagship for work a cheaper model passes.
Plans and pricing
5 free Nemotron (NVIDIA) models
$0 on OpenRouter free variantsFree variants such as nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free and nvidia/nemotron-3-super-120b-a12b:free. Free tiers are rate-limited and can change without notice.
Best for
Prototyping and benchmarking before you commit spend.
Not for
Production traffic that needs predictable throughput.
Nemotron 3.5 Lightning
$0.059 input / $0.17 output per 1M tokens262K context on OpenRouter (nvidia/nemotron-3.5-lightning). 1M input plus 1M output tokens: $0.23.
Best for
High-volume, low-stakes steps: routing, tagging, extraction and first drafts.
Not for
Tasks where a wrong answer costs more than the tokens saved.
Nemotron 3.5 Content Safety
$0.20 input / $0.20 output per 1M tokens131K context on OpenRouter (nvidia/nemotron-3.5-content-safety). 1M input plus 1M output tokens: $0.40.
Best for
A middle tier to test before paying flagship prices.
Not for
Tasks where a wrong answer costs more than the tokens saved.
Nemotron 3 Ultra
$0.50 input / $2.20 output per 1M tokens262K context on OpenRouter (nvidia/nemotron-3-ultra-550b-a55b). 1M input plus 1M output tokens: $2.70.
Best for
The hardest Nemotron (NVIDIA) tasks, once cheaper Nemotron (NVIDIA) models have failed your acceptance tests.
Not for
Bulk work a smaller model already passes; the price gap compounds at volume.
Prices are OpenRouter's listed per-token prices for NVIDIA models, read from the OpenRouter models API and checked 2026-10-04. Buying direct from NVIDIA or another host can be priced differently; prices exclude tax and change often.
Who should pay for Nemotron (NVIDIA)?
Nemotron (NVIDIA) suits builders who want free or near-free reasoning models to prototype and benchmark against paid ones. Start with Nemotron 3.5 Lightning on a sample of real tasks, measure the acceptance rate, and step up a tier only where it fails. Price the whole job, not the token: retries and human review often cost more than the model.
Calculate your exact cost →Who is overpaying for Nemotron (NVIDIA)?
- Teams sending every request to Nemotron 3 Ultra when Nemotron 3.5 Lightning passes most of them
- Anyone paying for output tokens they never read: cap max tokens and ask for concise answers
- Buyers comparing token prices without counting retries, failed responses and review time
Official pricing sources
Frequently asked questions
What is the cheapest Nemotron (NVIDIA) model?
On OpenRouter, Nemotron 3.5 Lightning at $0.059 input and $0.17 output per 1M tokens (checked 2026-10-04); 5 free variants are also listed, rate-limited.
How much does the top Nemotron (NVIDIA) model cost?
Nemotron 3 Ultra is listed at $0.50 input and $2.20 output per 1M tokens, so 1M of each costs $2.70.
Is Nemotron (NVIDIA) cheaper than GPT or Claude?
Usually per token, yes. Compare on your own prompts with the AI Cost Calculator: the cheaper model only saves money if it passes the same acceptance tests.