OverpayingForAIPricing desk

Classify · 5 checks · max 400 tokens · benched 2026-09-16

Kimi K3 vs Kimi K2.7 Code vs Kimi K2.6 vs Kimi K2.5 on sentiment of five reviews

Moonshotai models side by side on "Sentiment of five reviews": Kimi K2.5 scores 10/10; Kimi K2.5 is the cheapest answer scoring 8+ at $0.61 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a classifier. Output exactly the requested lines and nothing else.

User

Label each review as positive, neutral or negative. Output five lines in the form "<number>: <label>" and nothing else.

1. Battery lasts two full days and the screen is gorgeous. Would buy again.
2. It arrived on Tuesday in a brown box.
3. Stopped charging after three weeks and support never replied.
4. Does what it says. Nothing special, nothing wrong.
5. The strap broke on day one and the app crashes constantly.

Rubric for the judge: Five correctly labelled lines in the exact format; review 2 and 4 are neutral; no explanation.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Kimi K3

moonshotai/kimi-k3

10/10
Latency
1.2s
Cost
$0.00229
Per 1,000
$2.29

208 in · 157 out (124 reasoning) · 10 words · checks 5/5

  • regex
  • regex
  • regex
  • regex
  • max_words
1: positive
2: neutral
3: negative
4: neutral
5: negative

Judge: All five reviews correctly labeled in exact format with no extra commentary; reviews 2 and 4 properly identified as neutral.

Kimi K2.7 Code

moonshotai/kimi-k2.7-code

10/10
Latency
1.6s
Cost
$0.00070
Per 1,000
$0.70

119 in · 175 out (154 reasoning) · 10 words · checks 5/5

  • regex
  • regex
  • regex
  • regex
  • max_words
1: positive
2: neutral
3: negative
4: neutral
5: negative

Judge: All five reviews correctly labeled in exact format with no extra commentary; reviews 2 and 4 properly identified as neutral.

Kimi K2.6

moonshotai/kimi-k2.6

10/10
Latency
792ms
Cost
$0.00106
Per 1,000
$1.06

120 in · 236 out (217 reasoning) · 10 words · checks 5/5

  • regex
  • regex
  • regex
  • regex
  • max_words
1: positive
2: neutral
3: negative
4: neutral
5: negative

Judge: All five reviews correctly labeled in exact format with no extra commentary; reviews 2 and 4 properly identified as neutral.

Kimi K2.5

moonshotai/kimi-k2.5

10/10
top judge scorecheapest scoring 8+
Latency
2.2s
Cost
$0.00061
Per 1,000
$0.61

119 in · 249 out (228 reasoning) · 10 words · checks 5/5

  • regex
  • regex
  • regex
  • regex
  • max_words
1: positive
2: neutral
3: negative
4: neutral
5: negative

Judge: All five reviews correctly labeled in exact format with no extra commentary; reviews 2 and 4 properly identified as neutral.

Frequently asked

What does this prompt test?

Classify: Five correctly labelled lines in the exact format; review 2 and 4 are neutral; no explanation. The deterministic checks are regex, regex, regex, regex, max_words.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Kimi K2.5 at $0.61 per 1,000 runs. If you need the top score, Kimi K2.5 at $0.61 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.