OverpayingForAIPricing desk

Reasoning · 1 checks · max 800 tokens · benched 2026-09-16

Claude Opus 5 vs Claude Sonnet 5 vs Claude Sonnet 4.6 vs Claude Haiku 4.5 on count sellable units

Anthropic models side by side on "Count sellable units": Claude Sonnet 5 scores 10/10; Claude Sonnet 5 is the cheapest answer scoring 8+ at $0.82 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

Think carefully, then give only the final answer in the requested format.

User

A warehouse receives 3 pallets. Each pallet holds 48 boxes and each box holds 12 units. On inspection, 15% of the boxes are damaged; round the number of damaged boxes down to a whole box. The damaged boxes are discarded. Then 7 undamaged boxes are sent to the quality lab and are not sold. How many units are left to sell? Reply with only the number.

Rubric for the judge: The single correct answer 1392 (144 boxes, 21 damaged, 7 to the lab, 116 x 12).

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Claude Opus 5

anthropic/claude-opus-5

10/10
Latency
2.4s
Cost
$0.00192
Per 1,000
$1.92

154 in · 46 out (42 reasoning) · 1 words · checks 1/1

  • exact
1392

Judge: The output matches the correct answer of 1392 units exactly, provided in the required format of a bare number only.

Claude Sonnet 5

anthropic/claude-sonnet-5

10/10
top judge scorecheapest scoring 8+
Latency
979ms
Cost
$0.00082
Per 1,000
$0.82

154 in · 51 out (46 reasoning) · 1 words · checks 1/1

  • exact
1392

Judge: The output matches the correct answer of 1392 units exactly, provided in the required format of a bare number only.

Claude Sonnet 4.6

anthropic/claude-sonnet-4.6

10/10
Latency
759ms
Cost
$0.00179
Per 1,000
$1.79

112 in · 97 out (90 reasoning) · 1 words · checks 1/1

  • exact
1392

Judge: The output matches the correct answer of 1392 units exactly, provided in the required format of a bare number only.

Claude Haiku 4.5

anthropic/claude-haiku-4.5

10/10
Latency
851ms
Cost
$0.00085
Per 1,000
$0.85

111 in · 148 out · 68 words · checks 0/1

  • exact
Let me work through this step-by-step.

**Total boxes received:**
3 pallets × 48 boxes/pallet = 144 boxes

**Damaged boxes:**
15% of 144 = 0.15 × 144 = 21.6
Rounded down = 21 boxes

**Undamaged boxes after discarding damaged ones:**
144 - 21 = 123 boxes

**Boxes available to sell:**
123 - 7 = 116 boxes

**Units available to sell:**
116 boxes × 12 units/box = 1,392 units

1392

Judge: The output correctly calculates all steps (144 total boxes, 21 damaged, 123 undamaged, 116 available to sell, 1392 units) and provides only the required number as the final answer.

Frequently asked

What does this prompt test?

Reasoning: The single correct answer 1392 (144 boxes, 21 damaged, 7 to the lab, 116 x 12). The deterministic checks are exact.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Claude Sonnet 5 at $0.82 per 1,000 runs. If you need the top score, Claude Sonnet 5 at $0.82 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.