OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

Grok 4.6 vs Grok 4.5 vs Grok 4.20 vs Grok 4.3 on summarise a support thread in 60 words

xAI models side by side on "Summarise a support thread in 60 words": Grok 4.3 scores 9/10; Grok 4.3 is the cheapest answer scoring 8+ at $0.99 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Follow length limits exactly and output plain text only.

User

Summarise the following customer support thread in no more than 60 words. Mention what the customer wanted and how it was resolved.

---
Customer (Mon 09:12): Hi, I was charged twice for my Team plan on 1 September. Order references are 88213 and 88214. I only have one workspace. Can you fix this?
Agent (Mon 10:40): Sorry about that. I can see the duplicate charge. Could you confirm the last four digits of the card so I can raise the refund?
Customer (Mon 11:02): 4471.
Agent (Mon 11:30): Thanks. I have refunded order 88214 in full (USD 96.00). It should appear on your statement within 5 business days. I have also added a note so it cannot recur next cycle.
Customer (Mon 11:35): Great, thanks for the quick turnaround.
---

Rubric for the judge: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Grok 4.6

x-ai/grok-4.6

9/10
Latency
464ms
Cost
$0.00182
Per 1,000
$1.82

416 in · 196 out (147 reasoning) · 37 words · checks 3/3

  • max_words
  • contains
  • not_contains
The customer was charged twice for a Team plan and asked for a fix. The agent refunded duplicate order 88214 (USD 96.00), due on the statement within 5 business days, and noted the account to stop recurrence.

Judge: Accurate summary covering the issue (double charge), resolution (refund of $96), timeline (5 business days), and prevention measure (account note), within 60 words, with no card details leaked.

Grok 4.5

x-ai/grok-4.5

9/10
Latency
288ms
Cost
$0.00147
Per 1,000
$1.47

416 in · 143 out (94 reasoning) · 31 words · checks 3/3

  • max_words
  • contains
  • not_contains
Customer was double-charged for a Team plan (orders 88213 & 88214) despite one workspace. Agent confirmed, refunded 88214 in full (USD 96), and noted it to stop recurrence. Customer thanked them.

Judge: Accurate summary covering what customer wanted (fix double charge) and resolution (refund + prevention), within 60 words (31 words), no card digits leaked; minor: could explicitly state 5-day timeline but not critical.

Grok 4.20

x-ai/grok-4.20

9/10
Latency
242ms
Cost
$0.00189
Per 1,000
$1.89

393 in · 615 out (555 reasoning) · 45 words · checks 3/3

  • max_words
  • contains
  • not_contains
The customer reported a duplicate charge for their Team plan (orders 88213/88214) despite having one workspace and requested a refund. The agent verified details, issued a full USD 96 refund for order 88214, and added a note to prevent recurrence. The issue was resolved quickly.

Judge: Accurate, complete summary within 60 words (54 words), mentions what customer wanted and resolution, no card digits leaked, only minor issue is slight redundancy with 'quickly' at end.

Grok 4.3

x-ai/grok-4.3

9/10
top judge scorecheapest scoring 8+
Latency
295ms
Cost
$0.00099
Per 1,000
$0.99

399 in · 277 out (222 reasoning) · 39 words · checks 3/3

  • max_words
  • contains
  • not_contains
Customer reported duplicate Team plan charges (orders 88213/88214) and requested a refund. Agent confirmed the error, refunded USD 96 for order 88214 after card verification, and added a note to prevent recurrence; it should appear in 5 business days.

Judge: Accurate, complete summary covering the issue and resolution within 60 words (exactly 54), no card digits leaked, minor: could have explicitly stated 'one workspace' context but not required.

Frequently asked

What does this prompt test?

Summarise: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits. The deterministic checks are max_words, contains, not_contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Grok 4.3 at $0.99 per 1,000 runs. If you need the top score, Grok 4.3 at $0.99 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.