OverpayingForAIPricing desk

Nemotron (NVIDIA) Pricing: Which Model Is Worth Paying For?

OpenRouter lists 10 Nemotron (NVIDIA) models. For 1M input plus 1M output tokens they range from $0.23 (Nemotron 3.5 Lightning) to $2.70 (Nemotron 3 Ultra), a 12x spread. Most teams overpay by defaulting to the flagship for work a cheaper model passes.

LivePricing last verified: Oct 4, 2026Source: Official vendor pricing

Plans and pricing

5 free Nemotron (NVIDIA) models

$0 on OpenRouter free variants

Free variants such as nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free and nvidia/nemotron-3-super-120b-a12b:free. Free tiers are rate-limited and can change without notice.

Best for

Prototyping and benchmarking before you commit spend.

Not for

Production traffic that needs predictable throughput.

Nemotron 3.5 Lightning

$0.059 input / $0.17 output per 1M tokens

262K context on OpenRouter (nvidia/nemotron-3.5-lightning). 1M input plus 1M output tokens: $0.23.

Best for

High-volume, low-stakes steps: routing, tagging, extraction and first drafts.

Not for

Tasks where a wrong answer costs more than the tokens saved.

Nemotron 3.5 Content Safety

$0.20 input / $0.20 output per 1M tokens

131K context on OpenRouter (nvidia/nemotron-3.5-content-safety). 1M input plus 1M output tokens: $0.40.

Best for

A middle tier to test before paying flagship prices.

Not for

Tasks where a wrong answer costs more than the tokens saved.

Nemotron 3 Ultra

$0.50 input / $2.20 output per 1M tokens

262K context on OpenRouter (nvidia/nemotron-3-ultra-550b-a55b). 1M input plus 1M output tokens: $2.70.

Best for

The hardest Nemotron (NVIDIA) tasks, once cheaper Nemotron (NVIDIA) models have failed your acceptance tests.

Not for

Bulk work a smaller model already passes; the price gap compounds at volume.

Prices are OpenRouter's listed per-token prices for NVIDIA models, read from the OpenRouter models API and checked 2026-10-04. Buying direct from NVIDIA or another host can be priced differently; prices exclude tax and change often.

Who should pay for Nemotron (NVIDIA)?

Nemotron (NVIDIA) suits builders who want free or near-free reasoning models to prototype and benchmark against paid ones. Start with Nemotron 3.5 Lightning on a sample of real tasks, measure the acceptance rate, and step up a tier only where it fails. Price the whole job, not the token: retries and human review often cost more than the model.

Calculate your exact cost →

Who is overpaying for Nemotron (NVIDIA)?

  • Teams sending every request to Nemotron 3 Ultra when Nemotron 3.5 Lightning passes most of them
  • Anyone paying for output tokens they never read: cap max tokens and ask for concise answers
  • Buyers comparing token prices without counting retries, failed responses and review time

Official pricing sources

Frequently asked questions

What is the cheapest Nemotron (NVIDIA) model?

On OpenRouter, Nemotron 3.5 Lightning at $0.059 input and $0.17 output per 1M tokens (checked 2026-10-04); 5 free variants are also listed, rate-limited.

How much does the top Nemotron (NVIDIA) model cost?

Nemotron 3 Ultra is listed at $0.50 input and $2.20 output per 1M tokens, so 1M of each costs $2.70.

Is Nemotron (NVIDIA) cheaper than GPT or Claude?

Usually per token, yes. Compare on your own prompts with the AI Cost Calculator: the cheaper model only saves money if it passes the same acceptance tests.

Related pricing and comparisons

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

Pricing alerts

Want pricing changes before you overpay?

Get notified when AI plans, prices, or value-for-money signals change.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.