OverpayingForAIPricing desk

Lesson 2 of 8 · 10 min read · Beginner → Intermediate

Gemini models explained: Pro, Flash, Flash Lite and Gemma

The current Gemini model family, what each tier is for, how context windows work, and the live price ladder so you can see what a wrong choice costs.

In this lesson you will

  • Distinguish the Pro, Flash and Flash Lite tiers and explain the price gap between them.
  • Read the live price ladder and estimate the cost of a job on each model.
  • Know when a Gemma open-weight model is the cheaper answer.

Gemini is not one model. It is a ladder of models that trade quality for price, and the gap between rungs is large. Picking the right rung is the single biggest lever you have over cost. This lesson explains the rungs; the chart below shows what each one costs right now.

Current Gemini API models, $ per 1M tokensLive from the OverpayingForAI catalogue · last verified 2026-09-04 · sorted by output priceInputOutputGemini 3.1 Pro Preview1.0M context$2$12Gemini 2.5 Pro1.0M context$1.25$10Gemini 3.5 Flash1.0M context$1.5$9Gemini 3.6 Flash1.0M context$0.75$3.75Gemini 3.7 Flash1.0M context$0.75$3.75Gemini 3.8 Flash1.0M context$0.75$3.75Gemini 3 Flash Preview1.0M context$0.5$3Gemini 2.5 Flash1.0M context$0.3$2.5Batch, fast-mode and free variants are excluded. Prices change; the figure re-draws from the catalogue on every build.
Figure 1.Input and output prices per million tokens, pulled live from our catalogue. Output tokens are consistently several times more expensive than input, so verbose answers cost more than long prompts.

The three tiers

Pro is the flagship. It reasons more carefully, handles harder coding and analysis, and costs the most. At the time of writing Gemini 2.5 Pro is listed at $1.25 in and $10 out per million tokens, and Gemini 3.1 Pro Preview at $2 in and $12 out.

Flash is the workhorse. It is fast, cheap and good enough for most writing, summarising, extraction and simple code. The Flash line has several generations in the catalogue at once (2.5, 3, 3.5, 3.6, 3.7, 3.8) at prices from $0.30 to $1.50 input. Newer is not always cheaper — Gemini 3.5 Flash is listed above Gemini 3.8 Flash — so check the chart rather than assuming.

Flash Lite is the budget tier for high-volume, low-stakes work: classification, tagging, short rewrites, routing. Gemini 2.5 Flash Lite at $0.10 in and $0.40 out is the cheapest Gemini row we track. Gemini 3.1 Flash Lite and 3.5 Flash Lite sit a little higher.

Which tier for which job (prices at the time of writing, $ per 1M tokens)
TierExample modelInput / outputUse it forAvoid it for
ProGemini 3.1 Pro Preview$2.00 / $12.00Hard reasoning, complex code, long multi-step analysisBulk summarising, simple Q&A, anything Flash gets right
Pro (previous gen)Gemini 2.5 Pro$1.25 / $10.00Same as above when you want the cheaper ProLatency-sensitive chat
FlashGemini 3.8 Flash$0.75 / $3.75Everyday writing, summaries, extraction, moderate codeFrontier-level reasoning
FlashGemini 2.5 Flash$0.30 / $2.50Cheapest full Flash; great default for pipelinesLatest features
Flash LiteGemini 2.5 Flash Lite$0.10 / $0.40Classification, tagging, routing, short rewritesNuanced writing, multi-step reasoning
Open weightsGemma 4 31B$0.09 / $0.34Self-hosting, privacy, ultra-cheap bulk textLong context beyond 262K tokens, multimodal

Context window: big, but not free

Every current Gemini API model in our catalogue lists a 1,048,576-token context window. That is the headline feature. But context is not free: every token you put in the prompt is billed as input on every call. Pasting a 300-page PDF into a Pro prompt costs real money each time you ask a follow-up question. Lesson 7 explains context caching, which is the fix.

Gemma: the open-weight cousin

Gemma models (ids starting google/gemma-) are open weights. You can download them, run them on your own hardware, or rent them from hosts at very low prices — Gemma 4 31B is listed at $0.09 in and $0.34 out, and some hosts offer free-tier rows. They are text-first, have smaller context windows (131K to 262K tokens in our data) and lag Gemini on hard tasks. For simple, private or enormous-volume text jobs they can beat every Gemini price.

Knowledge check

You need to tag 200,000 support tickets by topic. Which model should you try first?

Lesson FAQ

Why are there so many Flash versions?

Google ships Flash frequently and keeps older versions available for developers who have tested against them. For a new project, pick the newest Flash that fits your budget on the chart; for an existing project, do not switch versions without re-testing.

What does "Preview" mean in a model name?

It is a model Google has released for use but may still change or retire. Fine for experiments; be cautious about building a production pipeline on one.

Is Gemini 1.5 still worth using?

Gemini 1.5 Flash and Pro still appear in our catalogue at very low prices, but Google has moved on. Use them only if you have existing code that depends on them.

Finished reading?

Mark it done to track your progress through the course.

Compare, calculate, decide — for Gemini

If our calculators helped you cut down on hidden AI wallet leaks, consider buying us a coffee. A tiny fraction of your savings keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.