OverpayingForAIPricing desk
9 min read·Last reviewed for accuracy · 2026-09-07·Prices verified · 2026-09-12

Fireworks Raised GPU Prices on 1 September. Watch for the Pass-Through

On-demand H100 hourly rates went from $7.00 to $8.00 and B200 from $10.00 to $13.00. Compute cost rises do not stay in the compute layer — they arrive later as inference prices.

The article text carries the review date. The rate table below is rebuilt from the live catalogue on every deploy.

Fastest win

If you rent GPUs directly, this is a bill you can see. If you buy inference by the token, it is a bill you cannot — yet. Compute price rises reach token prices eventually, and the teams who cope well are the ones who already know which of their workloads are price-sensitive and which are not.

What was published

Fireworks AI's pricing page, captured on our desk on 1 September 2026, states that on-demand GPU pricing increased from that date. H100 80GB and H200 141GB rise from $7.00 to $8.00 per hour. B200 goes from $10.00 to $13.00, B300 from $12.00 to $15.00, and GB300 from $18.00 to $20.00 per hour.

Those are increases of roughly 11% to 30% depending on the class, with the newest and most capable hardware taking the largest rises.

The same window recorded Google publishing detailed Vertex AI training and prediction rates — AutoML image training at $3.465/hour, tabular at $21.252/hour, Edge on-device at $18.00/hour — and noting that lower-cost machine types and scale-to-zero are no longer supported on the Agent Platform, with usage now charged in 30-second increments.

Why a GPU rate rise is a token-price signal

Almost every per-token price you pay is a compute cost with a margin on it. When the underlying hourly rate rises across the industry, per-token prices have a reason to follow.

The honest caveat: the lag is long and irregular. Competition, efficiency gains and new hardware generations all push the other way, and token prices have generally fallen through 2026 even as demand rose. One provider raising on-demand rates is not a forecast.

What it is, is a reminder that the direction is not guaranteed. Most AI budgets built over the last two years quietly assume per-token prices only go down. That assumption has already been challenged once this month, by Google publishing a dated doubling of Gemini Flash rates for January 2027.

Structural advice

Know which workloads survive a 30% rate rise, before one arrives

Nobody can tell you when a compute cost rise reaches your token bill. What you can do now is identify which pipelines are viable only at today's rate. Those are the ones to fix, and the exercise is worth doing whether or not the rise lands.

The removal of scale-to-zero is the sharper change

Buried in the Vertex capture is something that will cost some teams more than any hourly rate: lower-cost machine types and scale-to-zero are no longer supported on the Agent Platform.

Scale-to-zero is what makes intermittent workloads cheap. An endpoint that serves forty requests a day and costs nothing in between is a different economic proposition to one that must stay warm. Removing it does not change the hourly rate at all — it changes how many hours you are billed for.

If you have prototype or low-traffic endpoints on that platform, price them at always-on hours now. Some of them will no longer justify a dedicated endpoint, and the answer will be to batch the work or move it behind a shared service. Charging in 30-second increments with no minimum helps, but not enough to rescue a genuinely idle endpoint.

The stress test

List your top five AI workloads by monthly spend. For each one, write down what it costs today and what it would cost at a 30% higher rate — the top of the Fireworks range.

Now ask a single question of each: would we still run this? Not "could we afford it" — would we choose to. A workload that is obviously worth it at both prices needs no attention. A workload that only makes sense at today's rate is fragile, and you have just identified it for free.

For the fragile ones, the fixes are the ordinary ones and they are worth doing regardless: shorter context, capped output length, a cheaper model behind a quality gate, batch tiers where the vendor offers them, and committed capacity instead of on-demand where volume is predictable.

Committed capacity is a real option again

On-demand pricing is a convenience premium. When on-demand rates rise and your volume is predictable, the arithmetic behind a commitment improves without the commitment itself changing.

The test is utilisation. If you would use committed capacity above roughly 60% of the time, a commitment usually wins. Below that, on-demand flexibility is worth the premium and you should keep paying it.

Be honest about the forecast. The classic failure here is committing to capacity on the strength of a growth projection that does not arrive, and then paying for idle hardware for a year — a worse outcome than any hourly rate rise.

Ranked recommendation

Best move for everyone: run the 30% stress test this month. It costs an afternoon and it names your fragile workloads whether or not any rise reaches you.

Best move for predictable, high-utilisation workloads: price committed capacity against the new on-demand rates. The case is stronger this month than last for reasons that have nothing to do with your own usage.

Best move for intermittent endpoints: check whether scale-to-zero still exists on your platform, and re-price at always-on hours if it does not. This is the change most likely to surprise a research team's budget.

Avoid: restructuring your stack on one vendor's price rise, and any budget that assumes per-token prices only fall. Confirm current rates on each vendor's pricing page — these were captured on 1 September 2026.

Key Takeaways

  • Fireworks on-demand GPU rates rose from 1 September 2026: H100/H200 $7→$8, B200 $10→$13, B300 $12→$15, GB300 $18→$20 per hour
  • Vertex AI Agent Platform no longer supports scale-to-zero or lower-cost machine types — intermittent endpoints get dearer without any rate change
  • Compute rate rises reach per-token prices eventually, with a long and irregular lag
  • Stress-test your top five workloads at a 30% higher rate and find the ones that only work at today's price
  • Committed capacity is worth re-pricing when utilisation is above roughly 60%

Editorial context

Who is this for?

Research and ML teams who rent GPU capacity or run high-volume inference and are planning next year's budget.

When NOT to use this

Teams whose entire AI spend is a handful of subscription seats. Compute economics do not reach you in any actionable way.

Pricing insights

Hourly GPU rates and per-token inference rates are the same cost in different clothing. When one moves, the other has a reason to move, though the lag is long and irregular.

Alternatives to consider

Reserved or committed capacity instead of on-demand, smaller models where quality allows, batch tiers where offered, and reduced context and output length across the board.

Final verdict

Do not restructure on one vendor's price rise. Do run the stress test now, because the workloads that fail it will fail it whenever the rise comes.

Frequently Asked Questions

Will my token prices go up because of this?

Not necessarily, and not immediately. One provider raising on-demand GPU rates is a signal, not a forecast. Token prices have generally fallen through 2026. The useful response is a stress test, not a migration.

What is scale-to-zero and why does it matter?

It lets an endpoint cost nothing while idle. Without it, a low-traffic endpoint is billed for hours it does no work, which can cost more than an hourly rate rise. Check whether your platform still offers it.

Should we commit to capacity now?

Only if you would use it above roughly 60% of the time and your volume forecast is based on observed usage rather than a growth plan. An idle commitment is worse than an expensive on-demand hour.

Related

Free courses · no sign-up

Still deciding? Learn the basics first, then come back to the prices.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.