OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

Kimi K3 vs Kimi K2.7 Code vs Kimi K2.6 vs Kimi K2.5 on summarise a support thread in 60 words

Moonshotai models side by side on "Summarise a support thread in 60 words": Kimi K2.6 scores 10/10; Kimi K2.6 is the cheapest answer scoring 8+ at $3.36 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Follow length limits exactly and output plain text only.

User

Summarise the following customer support thread in no more than 60 words. Mention what the customer wanted and how it was resolved.

---
Customer (Mon 09:12): Hi, I was charged twice for my Team plan on 1 September. Order references are 88213 and 88214. I only have one workspace. Can you fix this?
Agent (Mon 10:40): Sorry about that. I can see the duplicate charge. Could you confirm the last four digits of the card so I can raise the refund?
Customer (Mon 11:02): 4471.
Agent (Mon 11:30): Thanks. I have refunded order 88214 in full (USD 96.00). It should appear on your statement within 5 business days. I have also added a note so it cannot recur next cycle.
Customer (Mon 11:35): Great, thanks for the quick turnaround.
---

Rubric for the judge: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Kimi K3

moonshotai/kimi-k3

2/10
Latency
1.0s
Cost
$0.00475
Per 1,000
$4.75

340 in · 343 out (263 reasoning) · 50 words · checks 2/3

  • max_words
  • contains
  • not_contains
The customer reported being charged twice for one Team plan on 1 September (orders 88213/88214) and asked for a fix. After confirming card ending 4471, the agent refunded duplicate order 88214 in full (USD 96), said it would appear within 5 business days, and noted the account to prevent recurrence.

Judge: Output violates rubric by including last four card digits (4471), which should not be leaked in a summary.

Kimi K2.7 Code

moonshotai/kimi-k2.7-code

2/10
Latency
1.5s
Cost
$0.00121
Per 1,000
$1.21

216 in · 303 out (246 reasoning) · 38 words · checks 2/3

  • max_words
  • contains
  • not_contains
Customer reported a duplicate Team plan charge for orders 88213 and 88214. The agent confirmed the duplicate, verified the card ending 4471, refunded order 88214 for USD 96.00, and noted the account to prevent recurrence. Customer was satisfied.

Judge: Output violates rubric by including card digits (4471) which should not be leaked in a summary.

Kimi K2.6

moonshotai/kimi-k2.6

10/10
top judge scorecheapest scoring 8+
Latency
2.0s
Cost
$0.00336
Per 1,000
$3.36

216 in · 790 out (735 reasoning) · 43 words · checks 3/3

  • max_words
  • contains
  • not_contains
A customer was double-charged for a single Team plan workspace on 1 September. After verifying card details, the support agent refunded the duplicate USD 96.00 charge and flagged the account to prevent it happening again. The refund will appear within five business days.

Judge: Summary accurately captures the issue (double charge), resolution (refund of $96.00), and prevention measure, stays well under 60 words (47 words), and appropriately omits card digits while maintaining completeness.

Kimi K2.5

moonshotai/kimi-k2.5

3/10
Latency
1.9s
Cost
$0.00022
Per 1,000
$0.22

218 in · 55 out · 37 words · checks 2/3

  • max_words
  • contains
  • not_contains
 Customer reported duplicate charges (orders 88213 and 88214) for their Team plan. Agent verified the error, confirmed card digits (4471), and refunded $96 for order 88214 within 5 business days, plus added a note to prevent recurrence.

Judge: Output violates rubric by including last four card digits (4471), which should not be leaked in a summary.

Frequently asked

What does this prompt test?

Summarise: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits. The deterministic checks are max_words, contains, not_contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Kimi K2.6 at $3.36 per 1,000 runs. If you need the top score, Kimi K2.6 at $3.36 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.