OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

Mistral Large 3 2512 vs Mistral Medium 3.5 vs Mistral Small 4 vs Ministral 3 8B 2512 on summarise a support thread in 60 words

Mistral AI models side by side on "Summarise a support thread in 60 words": Ministral 3 8B 2512 scores 9/10; Ministral 3 8B 2512 is the cheapest answer scoring 8+ at $0.05 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Follow length limits exactly and output plain text only.

User

Summarise the following customer support thread in no more than 60 words. Mention what the customer wanted and how it was resolved.

---
Customer (Mon 09:12): Hi, I was charged twice for my Team plan on 1 September. Order references are 88213 and 88214. I only have one workspace. Can you fix this?
Agent (Mon 10:40): Sorry about that. I can see the duplicate charge. Could you confirm the last four digits of the card so I can raise the refund?
Customer (Mon 11:02): 4471.
Agent (Mon 11:30): Thanks. I have refunded order 88214 in full (USD 96.00). It should appear on your statement within 5 business days. I have also added a note so it cannot recur next cycle.
Customer (Mon 11:35): Great, thanks for the quick turnaround.
---

Rubric for the judge: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Mistral Large 3 2512

mistralai/mistral-large-2512

9/10
Latency
479ms
Cost
$0.00021
Per 1,000
$0.21

238 in · 62 out · 35 words · checks 3/3

  • max_words
  • contains
  • not_contains
Customer reported a duplicate Team plan charge (orders 88213/88214). Agent verified, requested card details, then refunded the duplicate (USD 96) and flagged the account to prevent recurrence. Refund processed within 5 business days. Customer satisfied.

Judge: Accurate summary covering the issue, resolution, and outcome within 60 words (47 words), no card digits leaked, though could have been slightly more explicit about 'one workspace' context.

Mistral Medium 3.5

mistralai/mistral-medium-3-5

1/10
failed: empty output (finish_reason length)
Latency
1.2s
Cost
$0.01238
Per 1,000
$12.38

250 in · 1600 out (1282 reasoning) · 0 words · checks 2/3

  • max_words
  • contains
  • not_contains
(empty output)

Judge: Output is completely empty; fails to provide any summary despite clear task requirements.

Mistral Small 4

mistralai/mistral-small-2603

9/10
Latency
521ms
Cost
$0.00015
Per 1,000
$0.15

250 in · 189 out (163 reasoning) · 26 words · checks 3/3

  • max_words
  • contains
  • not_contains
Customer wanted refund for duplicate Team plan charge (USD 96.00). Agent confirmed, processed refund for order 88214, and prevented recurrence. Customer thanked agent for quick resolution.

Judge: Accurate summary covering what customer wanted (refund for duplicate charge) and resolution (refund processed, recurrence prevented), well under 60 words (27 words), no card digits leaked, minor deduction for not specifying the date or mentioning both order references.

Ministral 3 8B 2512

mistralai/ministral-8b-2512

9/10
top judge scorecheapest scoring 8+
Latency
318ms
Cost
$4.5e-5
Per 1,000
$0.05

238 in · 64 out · 28 words · checks 3/3

  • max_words
  • contains
  • not_contains
**Customer:** Reported duplicate Team plan charge (orders 88213, 88214) on 1 Sept.

**Resolution:** Agent refunded 88214 (USD 96) and blocked recurrence. Refund processed; appears in 5 business days.

Judge: Accurate summary covering the duplicate charge issue and complete resolution within 60 words (47 words), with no card details leaked; minor deduction for slightly informal formatting with bold headers rather than pure narrative.

Frequently asked

What does this prompt test?

Summarise: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits. The deterministic checks are max_words, contains, not_contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Ministral 3 8B 2512 at $0.05 per 1,000 runs. If you need the top score, Ministral 3 8B 2512 at $0.05 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.