OverpayingForAIPricing desk

Tool call · 5 checks · max 400 tokens · benched 2026-09-16

Kimi K3 vs Kimi K2.7 Code vs Kimi K2.6 vs Kimi K2.5 on pick the right tool of three

Moonshotai models side by side on "Pick the right tool of three": Kimi K2.7 Code scores 10/10; Kimi K2.7 Code is the cheapest answer scoring 8+ at $0.48 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a function-calling engine. Reply with JSON only, no prose, no code fences.

User

You have three tools:
1. {"name": "search_docs", "parameters": {"query": "string"}}
2. {"name": "create_ticket", "parameters": {"subject": "string", "body": "string"}}
3. {"name": "get_order_status", "parameters": {"order_id": "string"}}

Choose exactly one tool for the user message and produce the call as a JSON object of the shape {"name": string, "arguments": object}.

User: Where is my order A1029? It was supposed to arrive yesterday.

Rubric for the judge: Selects get_order_status with order_id A1029 and nothing else; does not open a ticket or search docs.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Kimi K3

moonshotai/kimi-k3

10/10
Latency
804ms
Cost
$0.00084
Per 1,000
$0.84

233 in · 60 out (32 reasoning) · 1 words · checks 5/5

  • json_valid
  • json_keys
  • regex
  • regex
  • not_contains
{"name":"get_order_status","arguments":{"order_id":"A1029"}}

Judge: Output correctly selects get_order_status with order_id A1029 in proper JSON format, matching the rubric requirements exactly.

Kimi K2.7 Code

moonshotai/kimi-k2.7-code

10/10
top judge scorecheapest scoring 8+
Latency
2.1s
Cost
$0.00048
Per 1,000
$0.48

147 in · 108 out (86 reasoning) · 5 words · checks 5/5

  • json_valid
  • json_keys
  • regex
  • regex
  • not_contains
{"name": "get_order_status", "arguments": {"order_id": "A1029"}}

Judge: Output correctly selects get_order_status with order_id A1029 in proper JSON format, matching the rubric requirements exactly.

Kimi K2.6

moonshotai/kimi-k2.6

10/10
Latency
594ms
Cost
$0.00072
Per 1,000
$0.72

148 in · 145 out (123 reasoning) · 5 words · checks 5/5

  • json_valid
  • json_keys
  • regex
  • regex
  • not_contains
 {"name": "get_order_status", "arguments": {"order_id": "A1029"}}

Judge: Output correctly selects get_order_status with order_id A1029 in proper JSON format, matching the rubric requirements exactly.

Kimi K2.5

moonshotai/kimi-k2.5

10/10
Latency
1.0s
Cost
$0.00049
Per 1,000
$0.49

148 in · 188 out (167 reasoning) · 5 words · checks 5/5

  • json_valid
  • json_keys
  • regex
  • regex
  • not_contains
{"name": "get_order_status", "arguments": {"order_id": "A1029"}}

Judge: Output correctly selects get_order_status with order_id A1029 in proper JSON format, matching the rubric requirements exactly.

Frequently asked

What does this prompt test?

Tool call: Selects get_order_status with order_id A1029 and nothing else; does not open a ticket or search docs. The deterministic checks are json_valid, json_keys, regex, regex, not_contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Kimi K2.7 Code at $0.48 per 1,000 runs. If you need the top score, Kimi K2.7 Code at $0.48 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.