OverpayingForAIPricing desk

Rewrite · 4 checks · max 400 tokens · benched 2026-09-16

Gemini 3.1 Pro Preview vs Gemini 3.8 Flash vs Gemini 3.5 Flash Lite on product update to a 280-character post

Google models side by side on "Product update to a 280-character post": Gemini 3.8 Flash scores 3/10. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

User

Turn this update into a single social media post of at most 280 characters. It must mention "v2.0" and end with the hashtag #OpenSource. Output only the post.

Update: Version 2.0 of our command-line tool is out. It is three times faster on large repositories, adds a built-in diff viewer, and is now licensed under MIT so anyone can contribute.

Rubric for the judge: Punchy post under 280 characters that mentions v2.0, the speed gain, the diff viewer and MIT, ending with #OpenSource.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

2/10
Latency
2.6s
Cost
$0.00492
Per 1,000
$4.92

86 in · 396 out (381 reasoning) · 8 words · checks 2/4

  • max_chars
  • contains
  • regex
  • contains
Our command-line tool v2.0 is officially out! 🚀

Judge: Output is only 56 characters and fails to mention speed gain, diff viewer, MIT license, or end with #OpenSource hashtag as required by rubric.

Gemini 3.8 Flash

google/gemini-3.8-flash

3/10
top judge score
Latency
1.3s
Cost
$0.00023
Per 1,000
$0.23

86 in · 44 out · 26 words · checks 4/4

  • max_chars
  • contains
  • regex
  • contains
v2.0 of our CLI tool is here! 🚀 

⚡ 3x faster on large repos
🔍 Built-in diff viewer
🤝 Now MIT-licensed so anyone can contribute

#OpenSource

Judge: Output exceeds 280 characters (312 total), contains extra commentary/emojis beyond what was requested, and fails the core constraint despite meeting content requirements.

Gemini 3.5 Flash Lite

google/gemini-3.5-flash-lite

2/10
Latency
1.2s
Cost
$0.00102
Per 1,000
$1.02

86 in · 396 out (384 reasoning) · 7 words · checks 2/4

  • max_chars
  • contains
  • regex
  • contains
CLI tool v2.0 is out! 🚀 Enjoy

Judge: Output is only 31 characters but fails to mention speed gain, diff viewer, MIT license, or end with #OpenSource hashtag as required by rubric.

Frequently asked

What does this prompt test?

Rewrite: Punchy post under 280 characters that mentions v2.0, the speed gain, the diff viewer and MIT, ending with #OpenSource. The deterministic checks are max_chars, contains, regex, contains.

Which model should I pick for this task?

Bench results are not in yet for this family and prompt.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.