OverpayingForAIPricing desk

Rewrite · 4 checks · max 400 tokens · benched 2026-09-16

GPT-5.5 vs GPT-5.4 vs GPT-5.4 mini vs GPT-5.4 Nano vs GPT-5.3-Codex on product update to a 280-character post

OpenAI models side by side on "Product update to a 280-character post": GPT-5.4 Nano scores 9/10; GPT-5.4 Nano is the cheapest answer scoring 8+ at $0.07 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

User

Turn this update into a single social media post of at most 280 characters. It must mention "v2.0" and end with the hashtag #OpenSource. Output only the post.

Update: Version 2.0 of our command-line tool is out. It is three times faster on large repositories, adds a built-in diff viewer, and is now licensed under MIT so anyone can contribute.

Rubric for the judge: Punchy post under 280 characters that mentions v2.0, the speed gain, the diff viewer and MIT, ending with #OpenSource.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

GPT-5.5

openai/gpt-5.5

9/10
Latency
349ms
Cost
$0.00250
Per 1,000
$2.50

87 in · 69 out (24 reasoning) · 27 words · checks 4/4

  • max_chars
  • contains
  • regex
  • contains
v2.0 is here! Our CLI tool is 3x faster on large repos, adds a built-in diff viewer, and is now MIT licensed so anyone can contribute. #OpenSource

Judge: Post is 139 characters, mentions v2.0, speed gain, diff viewer, MIT licensing, and ends with #OpenSource as required; only minor issue is it doesn't explicitly state 'anyone can contribute' though it's implied by 'MIT licensed so anyone can contribute.'

GPT-5.4

openai/gpt-5.4

9/10
Latency
392ms
Cost
$0.00122
Per 1,000
$1.22

87 in · 67 out (21 reasoning) · 27 words · checks 4/4

  • max_chars
  • contains
  • regex
  • contains
v2.0 of our command-line tool is here: 3x faster on large repositories, now with a built-in diff viewer, and licensed under MIT so anyone can contribute. #OpenSource

Judge: Post is 179 characters, mentions v2.0, speed gain (3x faster), diff viewer, and MIT licensing, ends with #OpenSource as required; only minor issue is it could be slightly punchier but meets all rubric criteria.

GPT-5.4 mini

openai/gpt-5.4-mini

9/10
Latency
357ms
Cost
$0.00037
Per 1,000
$0.37

87 in · 67 out (23 reasoning) · 23 words · checks 4/4

  • max_chars
  • contains
  • regex
  • contains
v2.0 of our command-line tool is here: 3x faster on large repos, built-in diff viewer, and now MIT-licensed so anyone can contribute. #OpenSource

Judge: Post is 139 characters, mentions v2.0, speed gain (3x faster), diff viewer, MIT licensing, and ends with #OpenSource as required; only minor issue is it doesn't explicitly state 'large repositories' but 'large repos' is a clear abbreviation.

GPT-5.4 Nano

openai/gpt-5.4-nano

9/10
top judge scorecheapest scoring 8+
Latency
256ms
Cost
$7.5e-5
Per 1,000
$0.07

87 in · 46 out · 25 words · checks 4/4

  • max_chars
  • contains
  • regex
  • contains
v2.0 is out! Our command-line tool is 3x faster on large repositories, includes a built-in diff viewer, and is now MIT-licensed—so everyone can contribute. #OpenSource

Judge: Post is 271 characters, mentions v2.0, speed gain, diff viewer, and MIT licensing, ends with #OpenSource as required; only minor issue is it could be slightly punchier but otherwise excellent.

GPT-5.3-Codex

openai/gpt-5.3-codex

9/10
Latency
394ms
Cost
$0.00207
Per 1,000
$2.07

87 in · 137 out (91 reasoning) · 27 words · checks 4/4

  • max_chars
  • contains
  • regex
  • contains
v2.0 of our command-line tool is here: 3x faster on large repos, now with a built-in diff viewer, and licensed under MIT so anyone can contribute. #OpenSource

Judge: Post is 179 characters, mentions v2.0, speed gain (3x faster), diff viewer, MIT license, and ends with #OpenSource as required; only minor issue is slightly less punchy tone than optimal.

Frequently asked

What does this prompt test?

Rewrite: Punchy post under 280 characters that mentions v2.0, the speed gain, the diff viewer and MIT, ending with #OpenSource. The deterministic checks are max_chars, contains, regex, contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, GPT-5.4 Nano at $0.07 per 1,000 runs. If you need the top score, GPT-5.4 Nano at $0.07 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.