OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

Qwen3.8 Max (0902) vs Qwen3.8 27B vs Qwen3.8 Flash vs Qwen3.7 Plus on summarise a support thread in 60 words

Alibaba models side by side on "Summarise a support thread in 60 words": Qwen3.7 Plus scores 9/10; Qwen3.7 Plus is the cheapest answer scoring 8+ at $1.40 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Follow length limits exactly and output plain text only.

User

Summarise the following customer support thread in no more than 60 words. Mention what the customer wanted and how it was resolved.

---
Customer (Mon 09:12): Hi, I was charged twice for my Team plan on 1 September. Order references are 88213 and 88214. I only have one workspace. Can you fix this?
Agent (Mon 10:40): Sorry about that. I can see the duplicate charge. Could you confirm the last four digits of the card so I can raise the refund?
Customer (Mon 11:02): 4471.
Agent (Mon 11:30): Thanks. I have refunded order 88214 in full (USD 96.00). It should appear on your statement within 5 business days. I have also added a note so it cannot recur next cycle.
Customer (Mon 11:35): Great, thanks for the quick turnaround.
---

Rubric for the judge: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Qwen3.8 Max (0902)

qwen/qwen3.8-max-0902

9/10
Latency
1.3s
Cost
$0.00154
Per 1,000
$1.54

277 in · 164 out (107 reasoning) · 41 words · checks 3/3

  • max_words
  • contains
  • not_contains
The customer reported a duplicate charge for their Team plan on 1 September. The agent confirmed the error, refunded the duplicate order (USD 96.00) in full, and added a note to prevent recurrence. The customer was satisfied with the quick resolution.

Judge: Accurate summary covering what the customer wanted (duplicate charge fixed) and resolution (refund issued, prevention note added), well under 60 words (38 words), no card digits leaked, minor deduction for not explicitly mentioning the 5-business-day timeline.

Qwen3.8 27B

qwen/qwen3.8-27b

2/10
Latency
772ms
Cost
$0.00108
Per 1,000
$1.08

295 in · 400 out (239 reasoning) · 19 words · checks 2/3

  • max_words
  • contains
  • not_contains


The customer reported a duplicate charge (USD 96.00) for their Team plan on 1 September, referencing orders 88213 and 

Judge: Output is incomplete and cuts off mid-sentence before explaining the resolution; fails to meet the 60-word requirement and doesn't clearly state how the issue was resolved.

Qwen3.8 Flash

qwen/qwen3.8-flash

7/10
Latency
1.1s
Cost
$0.00023
Per 1,000
$0.23

289 in · 400 out (316 reasoning) · 51 words · checks 3/3

  • max_words
  • contains
  • not_contains
Customer wanted a refund after being charged twice for their Team plan on 1 September (orders 88213 and 88214) despite having one workspace. Agent confirmed the duplicate, verified the card, refunded order 88214 in full (USD 96.00), said it would appear within 5 business days, and added a note to prevent

Judge: Summary is accurate and complete, stays under 60 words, avoids card digits, but is incomplete—it cuts off mid-sentence at 'prevent' without finishing the thought about preventing recurrence.

Qwen3.7 Plus

qwen/qwen3.7-plus

9/10
top judge scorecheapest scoring 8+
Latency
1.9s
Cost
$0.00140
Per 1,000
$1.40

251 in · 1032 out (990 reasoning) · 35 words · checks 3/3

  • max_words
  • contains
  • not_contains
The customer wanted a refund for a duplicate Team plan charge. The agent verified the card details, refunded the duplicate order in full, and added a note to prevent the charge from recurring next cycle.

Judge: Accurate summary covering what the customer wanted and how it was resolved, within 60 words (28 words), no card digits leaked, though minor detail omission of the refund amount and timeline.

Frequently asked

What does this prompt test?

Summarise: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits. The deterministic checks are max_words, contains, not_contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Qwen3.7 Plus at $1.40 per 1,000 runs. If you need the top score, Qwen3.7 Plus at $1.40 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.