OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

Grok 4.6 vs Grok 4.5 vs Grok 4.20 vs Grok 4.3 on meeting notes to exactly three bullets

xAI models side by side on "Meeting notes to exactly three bullets": Grok 4.3 scores 10/10; Grok 4.3 is the cheapest answer scoring 8+ at $1.53 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Output only what is asked for.

User

Turn these meeting notes into exactly three bullet points. Each bullet starts with "- " and is at most 20 words. Output nothing except the three bullets.

Notes: Launch of the new billing page moves from 18 Sept to 25 Sept because the tax calculation for EU customers still fails for VAT-exempt businesses. Priya owns the fix and will demo it Thursday. Marketing will hold the announcement email until the fix is verified. Separately, the team agreed to drop support for the legacy CSV export in Q1 next year and to notify affected customers in November.

Rubric for the judge: Three bullets that capture the delayed launch, the VAT fix owner, and the CSV export decision; nothing extra.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Grok 4.6

x-ai/grok-4.6

9/10
Latency
578ms
Cost
$0.00362
Per 1,000
$3.62

335 in · 523 out (468 reasoning) · 43 words · checks 3/3

  • regex
  • max_words
  • contains
- Billing page launch delayed to 25 Sept; Priya owns EU VAT-exempt tax fix, demos Thursday.
- Marketing holds announcement email until the tax calculation fix is verified.
- Drop legacy CSV export support in Q1 next year; notify affected customers in November.

Judge: All three bullets capture required points within 20 words each with no extra commentary, though the first bullet is slightly dense in packing information.

Grok 4.5

x-ai/grok-4.5

10/10
Latency
287ms
Cost
$0.00302
Per 1,000
$3.02

335 in · 428 out (379 reasoning) · 40 words · checks 3/3

  • regex
  • max_words
  • contains
- Billing page launch delayed to 25 Sept; Priya fixing EU VAT tax, demo Thursday.
- Marketing holds announcement email until the tax fix is verified.
- Drop legacy CSV export support in Q1 next year; notify customers in November.

Judge: Three bullets exactly, each under 20 words, capturing all three required points (delayed launch, VAT fix owner/demo, CSV export decision) with no extra commentary.

Grok 4.20

x-ai/grok-4.20

9/10
Latency
243ms
Cost
$0.00507
Per 1,000
$5.07

314 in · 1923 out (1881 reasoning) · 35 words · checks 3/3

  • regex
  • max_words
  • contains
- New billing launch delayed to 25 Sept due to EU VAT tax failure.
- Priya owns the fix and will demo it Thursday.
- Marketing holds announcement until verified; drop legacy CSV in Q1.

Judge: All three bullets capture required points within 20 words each, but the third bullet slightly conflates two separate decisions (announcement hold and CSV drop) that could be clearer as distinct items.

Grok 4.3

x-ai/grok-4.3

10/10
top judge scorecheapest scoring 8+
Latency
366ms
Cost
$0.00153
Per 1,000
$1.53

320 in · 534 out (488 reasoning) · 38 words · checks 3/3

  • regex
  • max_words
  • contains
- New billing page launch postponed to 25 Sept over EU VAT tax issue.
- Priya to demo fix Thursday; marketing delays email until verified.
- Drop legacy CSV export support in Q1 next year, notify in November.

Judge: Three bullets, each under 20 words, capturing all three required points (delayed launch, VAT fix owner, CSV decision) with no extra commentary.

Frequently asked

What does this prompt test?

Summarise: Three bullets that capture the delayed launch, the VAT fix owner, and the CSV export decision; nothing extra. The deterministic checks are regex, max_words, contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Grok 4.3 at $1.53 per 1,000 runs. If you need the top score, Grok 4.3 at $1.53 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.