OverpayingForAIPricing desk

Lesson 4 of 5 · 7 min read · Beginner

When the cheap model is enough, and when to pay for the flagship

Two real results from the bench of 2026-09-12: a support-thread summary where the cheapest Claude beat the flagship at a seventh of the cost, and arithmetic prompts where the cheapest Mistral got every answer wrong. Learn to tell the two cases apart.

In this lesson you will

  • Recognise a task the cheap model handles as well as the flagship
  • Recognise a task where the flagship's price is buying correctness
  • Judge a pair of real answers with cost attached

Two prompts from the same bench, two opposite lessons. Both use real answers, paraphrased only enough to fit on the page, with the real cost per run alongside.

Case one: summarise a support thread in 60 words

The prompt gives a support conversation about a double charge and asks for a summary of at most 60 words that says what the customer wanted and how it was resolved. The thread includes the last four digits of a card, and the rubric says a summary must not repeat them. Here are the cheapest and the most expensive Claude models on the bench of 2026-09-12.

Which would you trust?

Which answer should you trust more?

Case two: three arithmetic prompts

The bench also asks for bare numeric answers to short word problems: when a faster train catches a slower one, how many sellable units are left in a warehouse, and a price after stacked discounts and tax. Here is the cheapest and the most expensive Mistral model on the same bench.

Mistral family, bench of 2026-09-12. Ministral 3 8B is the cheapest row; Mistral Medium 3.5 the most expensive per 1,000 runs. Costs are the real cost of that single run.
PromptMinistral 3 8B answerCostMistral Medium 3.5 answerCostCorrect
Train catch-up (minutes)13$0.00001690$0.0048890
Sellable warehouse units10080$0.0000171392$0.002421392
Discount stack, then tax206.20$0.000016199.80$0.00163199.80

The cheap model was three hundred times cheaper and wrong three times out of three. A wrong number in an invoice or a report costs more than any model does, so here the flagship's price is buying the only thing that matters. Notice also that in the Anthropic family, Haiku 4.5 got all three numbers right but showed its working when the prompt said *only the number*; that is a format failure, fixable with a firmer prompt, and it is a different kind of miss from a wrong answer.

Lesson FAQ

Is the cheapest AI model good enough?

For summarising, rewriting, classification and extraction with clear rules, usually yes: on the bench of 2026-09-12 Claude Haiku 4.5 passed the support-thread summary where Opus 5 failed, at a seventh of the cost. For multi-step arithmetic and reasoning, often no: the cheapest Mistral got all three word problems wrong.

Why did the expensive model fail a simple summary?

It repeated the card's last four digits, which the rubric forbade. Stronger models are not immune to instruction misses; they are better on average, not on every run. Checks exist because averages do not send emails; individual answers do.

How do I know which case my task is?

Ask whether you would notice a wrong answer before it cost anything. If yes (you read it before using it), test the cheap model with checks. If no (the answer goes straight into a number or a decision), start with a stronger model and test downwards.

Finished reading?

Mark it done to track your progress through the course.

Save progress across devices

Get a private link that restores your lessons on any device. Email is optional and only used to send you the link.

Compare, calculate, decide

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.