OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

GPT-5.5 vs GPT-5.4 vs GPT-5.4 mini vs GPT-5.4 Nano vs GPT-5.3-Codex on summarise a support thread in 60 words

OpenAI models side by side on "Summarise a support thread in 60 words": GPT-5.4 Nano scores 9/10; GPT-5.4 Nano is the cheapest answer scoring 8+ at $0.14 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Follow length limits exactly and output plain text only.

User

Summarise the following customer support thread in no more than 60 words. Mention what the customer wanted and how it was resolved.

---
Customer (Mon 09:12): Hi, I was charged twice for my Team plan on 1 September. Order references are 88213 and 88214. I only have one workspace. Can you fix this?
Agent (Mon 10:40): Sorry about that. I can see the duplicate charge. Could you confirm the last four digits of the card so I can raise the refund?
Customer (Mon 11:02): 4471.
Agent (Mon 11:30): Thanks. I have refunded order 88214 in full (USD 96.00). It should appear on your statement within 5 business days. I have also added a note so it cannot recur next cycle.
Customer (Mon 11:35): Great, thanks for the quick turnaround.
---

Rubric for the judge: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

GPT-5.5

openai/gpt-5.5

9/10
Latency
396ms
Cost
$0.00297
Per 1,000
$2.97

215 in · 63 out · 47 words · checks 3/3

  • max_words
  • contains
  • not_contains
Customer was charged twice for one Team plan and requested the duplicate be fixed. After confirming the card’s last four digits, support refunded duplicate order 88214 for USD 96.00, advised it should appear within 5 business days, and added a note to prevent recurrence next billing cycle.

Judge: Accurate, complete summary covering what customer wanted and resolution within 60 words (56 words), no card digits leaked, minor: could have explicitly stated 'September 1st' but context is clear.

GPT-5.4

openai/gpt-5.4

9/10
Latency
315ms
Cost
$0.00162
Per 1,000
$1.62

215 in · 72 out · 52 words · checks 3/3

  • max_words
  • contains
  • not_contains
Customer reported a duplicate Team plan charge on 1 September for orders 88213 and 88214 and wanted it fixed. After confirming the last four card digits, support refunded order 88214 in full for USD 96.00, advised it would appear within 5 business days, and added a note to prevent it happening again.

Judge: Accurate, complete summary within 60 words (57 words) covering the issue and resolution without leaking card digits, though minor: could have been slightly more concise.

GPT-5.4 mini

openai/gpt-5.4-mini

9/10
Latency
413ms
Cost
$0.00049
Per 1,000
$0.49

215 in · 74 out (11 reasoning) · 45 words · checks 3/3

  • max_words
  • contains
  • not_contains
The customer wanted help with a duplicate Team plan charge for 1 September, citing two order references but only one workspace. Support confirmed the duplicate charge, refunded order 88214 in full ($96.00), and noted it to prevent recurrence. The customer confirmed the issue was resolved.

Judge: Accurate summary covering what customer wanted (duplicate charge fixed) and resolution (refund issued, prevention noted), within 60 words (54 words), no card digits leaked, minor: could have been slightly more explicit about the 5-day timeline.

GPT-5.4 Nano

openai/gpt-5.4-nano

9/10
top judge scorecheapest scoring 8+
Latency
246ms
Cost
$0.00014
Per 1,000
$0.14

215 in · 76 out · 53 words · checks 2/3

  • max_words
  • contains
  • not_contains
The customer reported being charged twice for their Team plan on 1 September (orders 88213 and 88214) despite having one workspace. The agent confirmed the duplicate charge and asked for the card’s last four digits. After the customer provided 4471, the agent refunded order 88214 (USD 96) and noted it to prevent recurrence.

Judge: Accurate, complete summary within 60 words (57 words) covering the issue and resolution, but the output includes the card digits '4471' which violates the 'no leaked card digits' rubric requirement.

GPT-5.3-Codex

openai/gpt-5.3-codex

2/10
Latency
333ms
Cost
$0.00150
Per 1,000
$1.50

215 in · 80 out · 54 words · checks 2/3

  • max_words
  • contains
  • not_contains
Customer reported being charged twice for one Team plan (orders 88213 and 88214) and asked for a fix. After verifying the card’s last four digits (4471), the agent confirmed a duplicate charge, refunded order 88214 in full (USD 96.00), advised it would appear within 5 business days, and added a note to prevent recurrence.

Judge: Output violates rubric by including customer's card digits (4471), which should never be leaked in a summary.

Frequently asked

What does this prompt test?

Summarise: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits. The deterministic checks are max_words, contains, not_contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, GPT-5.4 Nano at $0.14 per 1,000 runs. If you need the top score, GPT-5.4 Nano at $0.14 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.