OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

DeepSeek V4 Pro vs DeepSeek V4 Flash 0423 vs DeepSeek V4.1 Flash on summarise a support thread in 60 words

DeepSeek models side by side on "Summarise a support thread in 60 words": DeepSeek V4 Flash 0423 scores 9/10; DeepSeek V4 Flash 0423 is the cheapest answer scoring 8+ at $0.02 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Follow length limits exactly and output plain text only.

User

Summarise the following customer support thread in no more than 60 words. Mention what the customer wanted and how it was resolved.

---
Customer (Mon 09:12): Hi, I was charged twice for my Team plan on 1 September. Order references are 88213 and 88214. I only have one workspace. Can you fix this?
Agent (Mon 10:40): Sorry about that. I can see the duplicate charge. Could you confirm the last four digits of the card so I can raise the refund?
Customer (Mon 11:02): 4471.
Agent (Mon 11:30): Thanks. I have refunded order 88214 in full (USD 96.00). It should appear on your statement within 5 business days. I have also added a note so it cannot recur next cycle.
Customer (Mon 11:35): Great, thanks for the quick turnaround.
---

Rubric for the judge: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

DeepSeek V4 Pro

deepseek/deepseek-v4-pro

9/10
Latency
2.6s
Cost
$0.00069
Per 1,000
$0.69

209 in · 315 out (265 reasoning) · 37 words · checks 3/3

  • max_words
  • contains
  • not_contains
The customer reported a double charge for their Team plan and requested a fix. The duplicate order 88214 was refunded in full (USD 96.00), with a note added to prevent future recurrence. The issue was resolved promptly.

Judge: Accurate summary covering the problem (double charge), resolution (refund of $96.00), and preventive measure, within 60 words (31 words), with no card details leaked; only minor omission of the specific date (1 September) and order references.

DeepSeek V4 Flash 0423

deepseek/deepseek-v4-flash

9/10
top judge scorecheapest scoring 8+
Latency
831ms
Cost
$2.3e-5
Per 1,000
$0.02

209 in · 68 out (27 reasoning) · 30 words · checks 3/3

  • max_words
  • contains
  • not_contains
Customer reported a duplicate charge for their Team plan. The agent refunded order 88214 ($96) and added a note to prevent recurrence. The refund should appear within 5 business days.

Judge: Accurate summary covering the issue (duplicate charge), resolution (refund of $96 for order 88214), and timeline (5 business days), within 60 words, with no card details leaked; only minor omission is not explicitly stating the customer had one workspace.

DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash

2/10
Latency
100ms
Cost
$0.00021
Per 1,000
$0.21

235 in · 288 out (220 reasoning) · 45 words · checks 2/3

  • max_words
  • contains
  • not_contains
Customer reported being charged twice for a Team plan on 1 September (orders 88213 and 88214) despite having one workspace. After confirming card ending 4471, the agent refunded order 88214 in full (USD 96.00) within 5 business days and added a note to prevent recurrence.

Judge: Output violates rubric by including last four card digits (4471), which constitutes a security/privacy breach despite the rubric's explicit 'no leaked card digits' requirement.

Frequently asked

What does this prompt test?

Summarise: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits. The deterministic checks are max_words, contains, not_contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, DeepSeek V4 Flash 0423 at $0.02 per 1,000 runs. If you need the top score, DeepSeek V4 Flash 0423 at $0.02 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.