Pricing the Hermes model: three ways to access the same weights
The same Hermes 4 model can cost you nothing but electricity, a metered per-token rate, or a flat monthly subscription — depending entirely on how you choose to access it.
In this lesson you will
→Name the three access paths for the Hermes model and roughly what each costs
→Read this site's own verified per-token rate for Hermes 4 405B
→Use the fill-in exercise to estimate your own monthly cost
Three ways to run Hermes 4, and what each actually costs.
Access path
What you pay
Best for
Self-host (Ollama, vLLM, your own or a rented GPU)
No per-token charge — your cost is compute/electricity
Heavy, steady usage where you already have or can justify the hardware
Occasional or bursty use with no infrastructure to manage
Nous Portal subscription
Reported Free / Plus $20 / Super $100 / Ultra $200 per month
Bundled access plus Nous's own tool gateway and managed Hermes Cloud hosting
Figure 1.Every current Hermes model rate this site tracks, pulled live — the number that matters if you're calling the model through an API rather than self-hosting or subscribing to Nous Portal.
Your turn
Which access path fits your use case?
I expect to use Hermes for [USE_CASE], roughly [VOLUME] a month. Given that, [ACCESS_PATH] is the better fit because [REASON].
Hint: Low, bursty volume favors the metered API. Heavy, steady volume with existing hardware favors self-hosting. Wanting a bundled tool gateway and managed hosting favors the Portal subscription.
Knowledge check
You already own a spare machine with a capable GPU that otherwise sits unused, and you plan to run Hermes constantly for a personal project. Which access path is most likely to save you money?
Lesson FAQ
▸How much does the Hermes model cost?
It depends how you access it: free (plus your own compute) if you self-host, $1.00 input / $3.00 output per 1M tokens through the OpenRouter API for Hermes 4 405B per this site's catalogue, or a reported Free/$20/$100/$200-per-month Nous Portal subscription for bundled hosted access.
▸Is self-hosting Hermes actually free?
There's no per-token API charge, but it isn't free in an absolute sense — you're paying for the compute (your own hardware's electricity, or a rented GPU) instead. It only clearly beats a metered API when that hardware would otherwise sit idle.
Finished reading?
Mark it done to track your progress through the course.
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.