OverpayingForAIPricing desk

Playground · the testing desk

Run the same prompt across every model in a family. See what each answer costs.

24 predetermined prompts across 8 task types, benched weekly on 41 current models. Each cell shows the output, the checks it passed, a judge score, latency and the cost per 1,000 runs from the live catalogue. Weight the criteria yourself; the leaderboard re-ranks instantly.

Families

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version. Last bench 2026-09-16.

4 models · 24 prompts benched 2026-09-16

Anthropic

Opus, Sonnet and Haiku on the same eight tasks. See when Haiku is enough.

  • Claude Opus 5$5.00 / $25.00 per 1M
  • Claude Sonnet 5$2.00 / $10.00 per 1M
  • Claude Sonnet 4.6$3.00 / $15.00 per 1M
  • Claude Haiku 4.5$1.00 / $5.00 per 1M
Top by judge
Claude Sonnet 5 · 9.8/10
Cheapest listed
Claude Haiku 4.5

5 models · 24 prompts benched 2026-09-16

OpenAI

GPT-5.5 down to 5.4 nano and the Codex line, priced from the catalogue.

  • GPT-5.5$5.00 / $30.00 per 1M
  • GPT-5.4$2.50 / $15.00 per 1M
  • GPT-5.4 mini$0.75 / $4.50 per 1M
  • GPT-5.4 Nano$0.20 / $1.25 per 1M
  • GPT-5.3-Codex$1.75 / $14.00 per 1M
Top by judge
GPT-5.4 Nano · 9.8/10
Cheapest listed
GPT-5.4 Nano

3 models · 24 prompts benched 2026-09-16

Google

Gemini Pro, Flash and Flash Lite, with reasoning tokens counted in the cost.

  • Gemini 3.1 Pro Preview$2.00 / $12.00 per 1M
  • Gemini 3.8 Flash$0.75 / $3.75 per 1M
  • Gemini 3.5 Flash Lite$0.30 / $2.50 per 1M
Top by judge
Gemini 3.8 Flash · 8.8/10
Cheapest listed
Gemini 3.5 Flash Lite

3 models · 24 prompts benched 2026-09-16

DeepSeek

V4 Pro against the Flash rows that cost a few cents per thousand runs.

  • DeepSeek V4 Pro$1.60 / $3.20 per 1M
  • DeepSeek V4 Flash 0423$0.07 / $0.13 per 1M
  • DeepSeek V4.1 Flash$0.15 / $0.60 per 1M
Top by judge
DeepSeek V4 Pro · 9.6/10
Cheapest listed
DeepSeek V4 Flash 0423

4 models · 24 prompts benched 2026-09-16

xAI

Current-generation models from this provider on the same predetermined prompts.

  • Grok 4.6$2.00 / $6.00 per 1M
  • Grok 4.5$2.00 / $6.00 per 1M
  • Grok 4.20$1.25 / $2.50 per 1M
  • Grok 4.3$1.25 / $2.50 per 1M
Top by judge
Grok 4.5 · 9.7/10
Cheapest listed
Grok 4.20

4 models · 24 prompts benched 2026-09-16

Mistral AI

Current-generation models from this provider on the same predetermined prompts.

  • Mistral Large 3 2512$0.50 / $1.50 per 1M
  • Mistral Medium 3.5$1.50 / $7.50 per 1M
  • Mistral Small 4$0.15 / $0.60 per 1M
  • Ministral 3 8B 2512$0.15 / $0.15 per 1M
Top by judge
Mistral Small 4 · 9.8/10
Cheapest listed
Ministral 3 8B 2512

4 models · 24 prompts benched 2026-09-16

Meta

Current-generation models from this provider on the same predetermined prompts.

  • Muse Spark 1.3$1.25 / $4.25 per 1M
  • Muse Glimmer 30B$0.30 / $1.10 per 1M
  • Llama 4 Maverick$0.20 / $0.70 per 1M
  • Llama 4 Scout$0.10 / $0.30 per 1M
Top by judge
Muse Glimmer 30B · 9.4/10
Cheapest listed
Llama 4 Scout

4 models · 24 prompts benched 2026-09-16

Alibaba

Current-generation models from this provider on the same predetermined prompts.

  • Qwen3.8 Max (0902)$2.00 / $6.00 per 1M
  • Qwen3.8 27B$0.21 / $2.55 per 1M
  • Qwen3.8 Flash$0.15 / $0.47 per 1M
  • Qwen3.7 Plus$0.32 / $1.28 per 1M
Top by judge
Qwen3.8 Flash · 9.5/10
Cheapest listed
Qwen3.8 Flash

3 models · 24 prompts benched 2026-09-16

Z Ai

Current-generation models from this provider on the same predetermined prompts.

  • GLM 5.3$1.40 / $4.40 per 1M
  • GLM 5.2$0.60 / $2.00 per 1M
  • GLM 5.3 Flash$0.15 / $0.50 per 1M
Top by judge
GLM 5.3 · 9.5/10
Cheapest listed
GLM 5.3 Flash

4 models · 24 prompts benched 2026-09-16

Moonshotai

Current-generation models from this provider on the same predetermined prompts.

  • Kimi K3$2.65 / $13.28 per 1M
  • Kimi K2.7 Code$0.71 / $3.50 per 1M
  • Kimi K2.6$0.95 / $4.00 per 1M
  • Kimi K2.5$0.45 / $2.25 per 1M
Top by judge
Kimi K2.6 · 9.5/10
Cheapest listed
Kimi K2.5

3 models · 24 prompts benched 2026-09-16

Minimax

Current-generation models from this provider on the same predetermined prompts.

  • MiniMax M3$0.30 / $1.20 per 1M
  • MiniMax M2.7$0.30 / $1.20 per 1M
  • MiniMax M2.5$0.27 / $1.08 per 1M
Top by judge
MiniMax M3 · 9.8/10
Cheapest listed
MiniMax M2.5

Prompt catalogue

Short, unambiguous tasks with machine-checkable answers. Each links to its side-by-side page for the families that have been benched on it.

Summarise

Extract JSON

Classify

Rewrite

Code fix

SQL

Reasoning

Tool call

Questions before you read the numbers

How are the playground results produced?

Every week a script runs the same predetermined prompts through every current model in a family via OpenRouter at temperature 0, records the output, latency and token usage, prices it from the same models.json catalogue the rest of the site uses, runs deterministic checks, and asks Claude Haiku 4.5 to grade the output against a fixed rubric. The results are committed as JSON, so every visitor reads the same numbers for free.

Can I run my own prompt across a family?

Yes. The run page lets you paste your own OpenRouter key, which stays in your browser's localStorage and is sent only to openrouter.ai, and runs any prompt across a family at your own cost. Without a key you get three runs a day on the cheapest model per family and any :free row, on the site's budget.

What does 'cost per 1,000 runs' mean?

The mean cost of one run of that prompt on that model, multiplied by 1,000. It is the number to compare when you are choosing a model for a batch job: a $0.10 difference per thousand runs is noise, a $4 difference is a budget line.

Why is the judge Claude Haiku 4.5 and not the model being tested?

A fixed, cheap, capable judge keeps grades comparable across families and weeks. It grades against a written rubric, never against its own answer, and the deterministic checks (valid JSON, required clauses, exact numbers) do not depend on any judge at all.

Bench spend so far: $1.58 across 984 runs. Source: this page and the committed JSON in the repository.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.