Two prompts from the same bench, two opposite lessons. Both use real answers, paraphrased only enough to fit on the page, with the real cost per run alongside.
Case one: summarise a support thread in 60 words
The prompt gives a support conversation about a double charge and asks for a summary of at most 60 words that says what the customer wanted and how it was resolved. The thread includes the last four digits of a card, and the rubric says a summary must not repeat them. Here are the cheapest and the most expensive Claude models on the bench of 2026-09-12.
Which would you trust?
Which answer should you trust more?
Case two: three arithmetic prompts
The bench also asks for bare numeric answers to short word problems: when a faster train catches a slower one, how many sellable units are left in a warehouse, and a price after stacked discounts and tax. Here is the cheapest and the most expensive Mistral model on the same bench.
| Prompt | Ministral 3 8B answer | Cost | Mistral Medium 3.5 answer | Cost | Correct |
|---|---|---|---|---|---|
| Train catch-up (minutes) | 13 | $0.000016 | 90 | $0.00488 | 90 |
| Sellable warehouse units | 10080 | $0.000017 | 1392 | $0.00242 | 1392 |
| Discount stack, then tax | 206.20 | $0.000016 | 199.80 | $0.00163 | 199.80 |
The cheap model was three hundred times cheaper and wrong three times out of three. A wrong number in an invoice or a report costs more than any model does, so here the flagship's price is buying the only thing that matters. Notice also that in the Anthropic family, Haiku 4.5 got all three numbers right but showed its working when the prompt said *only the number*; that is a format failure, fixable with a firmer prompt, and it is a different kind of miss from a wrong answer.