OverpayingForAIPricing desk

Minimax · 3 models · 24 prompts · 72 runs · benched 2026-09-16

MiniMax M3 vs MiniMax M2.7 vs MiniMax M2.5 on 24 tasks: cost, checks, judge score

Same prompts, same temperature, same judge. MiniMax M3 leads on judge score; the cheapest row costs $0.33 per 1,000 runs. Move the sliders to rank by what you care about.

Leaderboard

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version. Last bench 2026-09-16.

Criteria weights
#ModelCompositeJudgeChecksLatencyWordsCost / runPer 1,000 runs$/1M in · out
1MiniMax M2.5minimax/minimax-m2.5 · 1 failed97.09.7496%1.1s17$0.00033$0.33$0.27 · $1.08
2MiniMax M3minimax/minimax-m385.69.7997%1.2s19$0.00035$0.35$0.30 · $1.20
3MiniMax M2.7minimax/minimax-m2.761.49.2590%2.1s17$0.00053$0.53$0.30 · $1.20

Per prompt

Open a prompt to read every model's output side by side, with the checks, the judge's reason and reader votes.

  1. Summarise

    Summarise a support thread in 60 words

    Top judge score
    MiniMax M2.5 · 9/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.19/1k
  2. Summarise

    Meeting notes to exactly three bullets

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.7 · $0.40/1k
  3. Summarise

    One-sentence changelog summary

    Top judge score
    MiniMax M2.5 · 9/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.42/1k
  4. Extract JSON

    Invoice text to JSON

    Top judge score
    MiniMax M2.5 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.22/1k
  5. Extract JSON

    People mentioned to a JSON array

    Top judge score
    MiniMax M3 · 9/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.22/1k
  6. Extract JSON

    Event announcement to structured JSON

    Top judge score
    MiniMax M2.5 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.21/1k
  7. Classify

    Label six support tickets

    Top judge score
    MiniMax M2.7 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.7 · $0.43/1k
  8. Classify

    Sentiment of five reviews

    Top judge score
    MiniMax M2.5 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.26/1k
  9. Classify

    Single message intent

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.09/1k
  10. Rewrite

    Rewrite a stiff notice in a friendly tone

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.56/1k
  11. Rewrite

    Jargon to plain English for a 12-year-old

    Top judge score
    MiniMax M2.5 · 9/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.40/1k
  12. Rewrite

    Product update to a 280-character post

    Top judge score
    MiniMax M3 · 9/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.36/1k
  13. Code fix

    Fix an off-by-one in Python

    Top judge score
    MiniMax M2.5 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.16/1k
  14. Code fix

    Fix a missing await in JavaScript

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.14/1k
  15. Code fix

    Parameterise a SQL query in Python

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.12/1k
  16. SQL

    Top five customers by 2025 spend

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.27/1k
  17. SQL

    Monthly active users

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.40/1k
  18. SQL

    Products over 100 units with HAVING

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.13/1k
  19. Reasoning

    When does the faster train catch up?

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.38/1k
  20. Reasoning

    Count sellable units

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.20/1k
  21. Reasoning

    Stacked discounts and tax

    Top judge score
    MiniMax M2.5 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.25/1k
  22. Tool call

    Emit a weather tool call

    Top judge score
    MiniMax M2.7 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.7 · $0.09/1k
  23. Tool call

    Emit a calendar tool call

    Top judge score
    MiniMax M2.5 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M2.5 · $0.20/1k
  24. Tool call

    Pick the right tool of three

    Top judge score
    MiniMax M3 · 10/10
    Cheapest scoring ≥ 8
    MiniMax M3 · $0.12/1k

Frequently asked

Which Minimax model is cheapest per run on these tasks?

MiniMax M2.5 at $0.33 per 1,000 runs, measured from real token usage, not list price.

How is the composite score calculated?

Each model gets a 0-1 score on cost, quality (mean judge score), speed, checks passed and brevity, relative to the other models in the family. Your slider weights combine them into a 0-100 composite. Cost, speed and brevity are log-scaled because prices span decades.

Are these the official prices?

They are the catalogue rates in models.json, synced hourly from OpenRouter's pricing API, multiplied by the tokens each run actually used. The $/1M columns show the rate; the cost columns show what the run cost.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.