OverpayingForAIPricing desk

Lesson 2 of 3 · 7 min read · Beginner

Pricing the Hermes model: three ways to access the same weights

The same Hermes 4 model can cost you nothing but electricity, a metered per-token rate, or a flat monthly subscription — depending entirely on how you choose to access it.

In this lesson you will

  • Name the three access paths for the Hermes model and roughly what each costs
  • Read this site's own verified per-token rate for Hermes 4 405B
  • Use the fill-in exercise to estimate your own monthly cost
Three ways to run Hermes 4, and what each actually costs.
Access pathWhat you payBest for
Self-host (Ollama, vLLM, your own or a rented GPU)No per-token charge — your cost is compute/electricityHeavy, steady usage where you already have or can justify the hardware
OpenRouter API (verified, this site's catalogue)$1.00 input / $3.00 output per 1M tokens (Hermes 4 405B)Occasional or bursty use with no infrastructure to manage
Nous Portal subscriptionReported Free / Plus $20 / Super $100 / Ultra $200 per monthBundled access plus Nous's own tool gateway and managed Hermes Cloud hosting
Hermes model pricing, live from this site's catalogueLive from the OverpayingForAI catalogue · last verified 2026-09-12 · sorted by output priceInputOutputHermes 4 405B131K context$1$3Hermes 3 405B Instruct131K context$1$1Hermes 3 70B Instruct131K context$0.7$0.7Hermes 4 70B131K context$0.13$0.4Batch, fast-mode and free variants are excluded. Prices change; the figure re-draws from the catalogue on every build.
Figure 1.Every current Hermes model rate this site tracks, pulled live — the number that matters if you're calling the model through an API rather than self-hosting or subscribing to Nous Portal.

Your turn

Which access path fits your use case?

I expect to use Hermes for [USE_CASE], roughly [VOLUME] a month. Given that, [ACCESS_PATH] is the better fit because [REASON].

Hint: Low, bursty volume favors the metered API. Heavy, steady volume with existing hardware favors self-hosting. Wanting a bundled tool gateway and managed hosting favors the Portal subscription.

Knowledge check

You already own a spare machine with a capable GPU that otherwise sits unused, and you plan to run Hermes constantly for a personal project. Which access path is most likely to save you money?

Lesson FAQ

How much does the Hermes model cost?

It depends how you access it: free (plus your own compute) if you self-host, $1.00 input / $3.00 output per 1M tokens through the OpenRouter API for Hermes 4 405B per this site's catalogue, or a reported Free/$20/$100/$200-per-month Nous Portal subscription for bundled hosted access.

Is self-hosting Hermes actually free?

There's no per-token API charge, but it isn't free in an absolute sense — you're paying for the compute (your own hardware's electricity, or a rented GPU) instead. It only clearly beats a metered API when that hardware would otherwise sit idle.

Finished reading?

Mark it done to track your progress through the course.

Save progress across devices

Get a private link that restores your lessons on any device. Email is optional and only used to send you the link.

Compare, calculate, decide

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.