OverpayingForAIPricing desk

Rewrite · 6 checks · max 400 tokens · benched 2026-09-16

Gemini 3.1 Pro Preview vs Gemini 3.8 Flash vs Gemini 3.5 Flash Lite on jargon to plain english for a 12-year-old

Google models side by side on "Jargon to plain English for a 12-year-old": Gemini 3.5 Flash Lite scores 2/10. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

User

Rewrite this paragraph so a 12-year-old would understand it. Keep it under 70 words and do not use the words "leverage", "synergy" or "paradigm". Output only the rewritten paragraph.

Paragraph: Our next-generation battery architecture leverages a solid-state electrolyte paradigm to achieve synergy between energy density and thermal stability, enabling a 40% increase in range while mitigating the combustion risk profile associated with legacy lithium-ion chemistries.

Rubric for the judge: Simple vocabulary, keeps the 40% range gain and the safety point, under 70 words, banned words absent.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

2/10
Latency
2.4s
Cost
$0.00495
Per 1,000
$4.95

98 in · 396 out (383 reasoning) · 12 words · checks 5/6

  • max_words
  • contains
  • contains
  • not_contains
  • not_contains
  • not_contains
Our new battery design uses solid materials instead of liquids. This helps

Judge: Output is incomplete (cuts off mid-sentence), fails to mention the 40% range increase or safety benefits, and doesn't meet the 70-word requirement despite being unfinished.

Gemini 3.8 Flash

google/gemini-3.8-flash

2/10
Latency
2.7s
Cost
$0.00156
Per 1,000
$1.56

98 in · 396 out (380 reasoning) · 14 words · checks 5/6

  • max_words
  • contains
  • contains
  • not_contains
  • not_contains
  • not_contains
We designed a brand-new type of battery that uses solid materials instead of liquids

Judge: Output is only 16 words (well under 70), avoids banned words, but fails to mention the critical 40% range increase and safety benefits that the rubric explicitly requires.

Gemini 3.5 Flash Lite

google/gemini-3.5-flash-lite

2/10
top judge score
Latency
940ms
Cost
$0.00102
Per 1,000
$1.02

98 in · 396 out (381 reasoning) · 14 words · checks 5/6

  • max_words
  • contains
  • contains
  • not_contains
  • not_contains
  • not_contains
Our new battery design uses solid materials instead of liquid. This lets it hold

Judge: Output is incomplete (cuts off mid-sentence), fails to mention the 40% range increase or safety benefits, and doesn't meet the 70-word requirement despite being unfinished.

Frequently asked

What does this prompt test?

Rewrite: Simple vocabulary, keeps the 40% range gain and the safety point, under 70 words, banned words absent. The deterministic checks are max_words, contains, contains, not_contains, not_contains, not_contains.

Which model should I pick for this task?

Bench results are not in yet for this family and prompt.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.