OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

Claude Opus 5 vs Claude Sonnet 5 vs Claude Sonnet 4.6 vs Claude Haiku 4.5 on meeting notes to exactly three bullets

Anthropic models side by side on "Meeting notes to exactly three bullets": Claude Sonnet 5 scores 10/10; Claude Haiku 4.5 is the cheapest answer scoring 8+ at $0.52 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Output only what is asked for.

User

Turn these meeting notes into exactly three bullet points. Each bullet starts with "- " and is at most 20 words. Output nothing except the three bullets.

Notes: Launch of the new billing page moves from 18 Sept to 25 Sept because the tax calculation for EU customers still fails for VAT-exempt businesses. Priya owns the fix and will demo it Thursday. Marketing will hold the announcement email until the fix is verified. Separately, the team agreed to drop support for the legacy CSV export in Q1 next year and to notify affected customers in November.

Rubric for the judge: Three bullets that capture the delayed launch, the VAT fix owner, and the CSV export decision; nothing extra.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Claude Opus 5

anthropic/claude-opus-5

10/10
Latency
1.5s
Cost
$0.00348
Per 1,000
$3.48

211 in · 97 out · 42 words · checks 3/3

  • regex
  • max_words
  • contains
- Billing page launch moves to 25 Sept; EU VAT-exempt tax calculation still failing.
- Priya owns the fix and demos Thursday; marketing holds announcement email until verified.
- Legacy CSV export support ends Q1 next year; affected customers notified in November.

Judge: Three bullets, each under 20 words, capturing all three required points (delayed launch, VAT fix owner/demo, CSV export decision) with no extra commentary.

Claude Sonnet 5

anthropic/claude-sonnet-5

10/10
top judge score
Latency
1.3s
Cost
$0.00142
Per 1,000
$1.42

211 in · 100 out · 42 words · checks 3/3

  • regex
  • max_words
  • contains
- Billing page launch delayed from 18 Sept to 25 Sept due to EU VAT-exempt tax bug.
- Priya owns the fix, demoing Thursday; marketing holds announcement until verified.
- Legacy CSV export support drops in Q1; affected customers notified in November.

Judge: Three bullets exactly, each under 20 words, capturing all three required points (delayed launch, VAT fix owner, CSV decision) with no extra commentary.

Claude Sonnet 4.6

anthropic/claude-sonnet-4.6

9/10
Latency
1.2s
Cost
$0.00139
Per 1,000
$1.39

143 in · 64 out · 44 words · checks 3/3

  • regex
  • max_words
  • contains
- Launch delayed to 25 Sept due to VAT-exempt EU tax bug; Priya owns fix, demos Thursday.
- Marketing will hold announcement email until the billing page fix is verified.
- Legacy CSV export drops in Q1; affected customers will be notified in November.

Judge: All three bullets capture required points within 20-word limit with proper formatting, though the first bullet is slightly dense and could be clearer.

Claude Haiku 4.5

anthropic/claude-haiku-4.5

9/10
cheapest scoring 8+
Latency
710ms
Cost
$0.00052
Per 1,000
$0.52

142 in · 76 out · 53 words · checks 2/3

  • regex
  • max_words
  • contains
- Billing page launch delayed from Sept 18 to Sept 25 due to EU VAT-exempt tax calculation issues; Priya demos fix Thursday.
- Marketing holding announcement email until tax calculation fix is verified and confirmed working properly.
- Legacy CSV export support ending Q1 next year; affected customers will be notified in November.

Judge: All three bullets capture required points (delayed launch, VAT fix owner/demo, CSV export decision) with no extra commentary, though first bullet is 21 words (one over limit).

Frequently asked

What does this prompt test?

Summarise: Three bullets that capture the delayed launch, the VAT fix owner, and the CSV export decision; nothing extra. The deterministic checks are regex, max_words, contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Claude Haiku 4.5 at $0.52 per 1,000 runs. If you need the top score, Claude Sonnet 5 at $1.42 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.