Families
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version. Last bench 2026-09-16.
4 models · 24 prompts benched 2026-09-16
Opus, Sonnet and Haiku on the same eight tasks. See when Haiku is enough.
- Claude Opus 5$5.00 / $25.00 per 1M
- Claude Sonnet 5$2.00 / $10.00 per 1M
- Claude Sonnet 4.6$3.00 / $15.00 per 1M
- Claude Haiku 4.5$1.00 / $5.00 per 1M
- Top by judge
- Claude Sonnet 5 · 9.8/10
- Cheapest listed
- Claude Haiku 4.5
5 models · 24 prompts benched 2026-09-16
GPT-5.5 down to 5.4 nano and the Codex line, priced from the catalogue.
- GPT-5.5$5.00 / $30.00 per 1M
- GPT-5.4$2.50 / $15.00 per 1M
- GPT-5.4 mini$0.75 / $4.50 per 1M
- GPT-5.4 Nano$0.20 / $1.25 per 1M
- GPT-5.3-Codex$1.75 / $14.00 per 1M
- Top by judge
- GPT-5.4 Nano · 9.8/10
- Cheapest listed
- GPT-5.4 Nano
3 models · 24 prompts benched 2026-09-16
Gemini Pro, Flash and Flash Lite, with reasoning tokens counted in the cost.
- Gemini 3.1 Pro Preview$2.00 / $12.00 per 1M
- Gemini 3.8 Flash$0.75 / $3.75 per 1M
- Gemini 3.5 Flash Lite$0.30 / $2.50 per 1M
- Top by judge
- Gemini 3.8 Flash · 8.8/10
- Cheapest listed
- Gemini 3.5 Flash Lite
3 models · 24 prompts benched 2026-09-16
V4 Pro against the Flash rows that cost a few cents per thousand runs.
- DeepSeek V4 Pro$1.60 / $3.20 per 1M
- DeepSeek V4 Flash 0423$0.07 / $0.13 per 1M
- DeepSeek V4.1 Flash$0.15 / $0.60 per 1M
- Top by judge
- DeepSeek V4 Pro · 9.6/10
- Cheapest listed
- DeepSeek V4 Flash 0423
4 models · 24 prompts benched 2026-09-16
Current-generation models from this provider on the same predetermined prompts.
- Grok 4.6$2.00 / $6.00 per 1M
- Grok 4.5$2.00 / $6.00 per 1M
- Grok 4.20$1.25 / $2.50 per 1M
- Grok 4.3$1.25 / $2.50 per 1M
- Top by judge
- Grok 4.5 · 9.7/10
- Cheapest listed
- Grok 4.20
4 models · 24 prompts benched 2026-09-16
Current-generation models from this provider on the same predetermined prompts.
- Mistral Large 3 2512$0.50 / $1.50 per 1M
- Mistral Medium 3.5$1.50 / $7.50 per 1M
- Mistral Small 4$0.15 / $0.60 per 1M
- Ministral 3 8B 2512$0.15 / $0.15 per 1M
- Top by judge
- Mistral Small 4 · 9.8/10
- Cheapest listed
- Ministral 3 8B 2512
4 models · 24 prompts benched 2026-09-16
Current-generation models from this provider on the same predetermined prompts.
- Muse Spark 1.3$1.25 / $4.25 per 1M
- Muse Glimmer 30B$0.30 / $1.10 per 1M
- Llama 4 Maverick$0.20 / $0.70 per 1M
- Llama 4 Scout$0.10 / $0.30 per 1M
- Top by judge
- Muse Glimmer 30B · 9.4/10
- Cheapest listed
- Llama 4 Scout
4 models · 24 prompts benched 2026-09-16
Current-generation models from this provider on the same predetermined prompts.
- Qwen3.8 Max (0902)$2.00 / $6.00 per 1M
- Qwen3.8 27B$0.21 / $2.55 per 1M
- Qwen3.8 Flash$0.15 / $0.47 per 1M
- Qwen3.7 Plus$0.32 / $1.28 per 1M
- Top by judge
- Qwen3.8 Flash · 9.5/10
- Cheapest listed
- Qwen3.8 Flash
3 models · 24 prompts benched 2026-09-16
Current-generation models from this provider on the same predetermined prompts.
- GLM 5.3$1.40 / $4.40 per 1M
- GLM 5.2$0.60 / $2.00 per 1M
- GLM 5.3 Flash$0.15 / $0.50 per 1M
- Top by judge
- GLM 5.3 · 9.5/10
- Cheapest listed
- GLM 5.3 Flash
4 models · 24 prompts benched 2026-09-16
Current-generation models from this provider on the same predetermined prompts.
- Kimi K3$2.65 / $13.28 per 1M
- Kimi K2.7 Code$0.71 / $3.50 per 1M
- Kimi K2.6$0.95 / $4.00 per 1M
- Kimi K2.5$0.45 / $2.25 per 1M
- Top by judge
- Kimi K2.6 · 9.5/10
- Cheapest listed
- Kimi K2.5
3 models · 24 prompts benched 2026-09-16
Current-generation models from this provider on the same predetermined prompts.
- MiniMax M3$0.30 / $1.20 per 1M
- MiniMax M2.7$0.30 / $1.20 per 1M
- MiniMax M2.5$0.27 / $1.08 per 1M
- Top by judge
- MiniMax M3 · 9.8/10
- Cheapest listed
- MiniMax M2.5