DeepSeek · 3 models · 24 prompts · 72 runs · benched 2026-09-16
DeepSeek V4 Pro vs DeepSeek V4 Flash 0423 vs DeepSeek V4.1 Flash on 24 tasks: cost, checks, judge score
Same prompts, same temperature, same judge. DeepSeek V4 Pro leads on judge score; the cheapest row costs $0.10 per 1,000 runs. Move the sliders to rank by what you care about.
Leaderboard
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version. Last bench 2026-09-16.
| # | Model | Composite | Judge | Checks | Latency | Words | Cost / run | Per 1,000 runs | $/1M in · out |
|---|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | 96.0 | 9.46 | 92% | 832ms | 17 | $9.5e-5 | $0.10 | $0.07 · $0.13 |
| 2 | DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | 79.4 | 9.00 | 93% | 1.6s | 17 | $0.00014 | $0.14 | $0.15 · $0.60 |
| 3 | DeepSeek V4 Prodeepseek/deepseek-v4-pro | 55.7 | 9.63 | 95% | 1.7s | 18 | $0.00061 | $0.61 | $1.60 · $3.20 |
Per prompt
Open a prompt to read every model's output side by side, with the checks, the judge's reason and reader votes.
Summarise
Summarise a support thread in 60 words
- Top judge score
- DeepSeek V4 Flash 0423 · 9/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.02/1k
Summarise
Meeting notes to exactly three bullets
- Top judge score
- DeepSeek V4.1 Flash · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4.1 Flash · $0.26/1k
Summarise
One-sentence changelog summary
- Top judge score
- DeepSeek V4 Flash 0423 · 9/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.05/1k
Extract JSON
Invoice text to JSON
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.03/1k
Extract JSON
People mentioned to a JSON array
- Top judge score
- DeepSeek V4 Flash 0423 · 9/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.04/1k
Extract JSON
Event announcement to structured JSON
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.03/1k
Classify
Label six support tickets
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.04/1k
Classify
Sentiment of five reviews
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.02/1k
Classify
Single message intent
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.01/1k
Rewrite
Rewrite a stiff notice in a friendly tone
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.03/1k
Rewrite
Jargon to plain English for a 12-year-old
- Top judge score
- DeepSeek V4 Flash 0423 · 9/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.03/1k
Rewrite
Product update to a 280-character post
- Top judge score
- DeepSeek V4.1 Flash · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4.1 Flash · $0.19/1k
Code fix
Fix an off-by-one in Python
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.02/1k
Code fix
Fix a missing await in JavaScript
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.03/1k
Code fix
Parameterise a SQL query in Python
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.03/1k
SQL
Top five customers by 2025 spend
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.03/1k
SQL
Monthly active users
- Top judge score
- DeepSeek V4.1 Flash · 7/10
- Cheapest scoring ≥ 8
- none
SQL
Products over 100 units with HAVING
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.04/1k
Reasoning
When does the faster train catch up?
- Top judge score
- DeepSeek V4.1 Flash · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4.1 Flash · $0.06/1k
Reasoning
Count sellable units
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.05/1k
Reasoning
Stacked discounts and tax
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.02/1k
Tool call
Emit a weather tool call
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.02/1k
Tool call
Emit a calendar tool call
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.03/1k
Tool call
Pick the right tool of three
- Top judge score
- DeepSeek V4 Flash 0423 · 10/10
- Cheapest scoring ≥ 8
- DeepSeek V4 Flash 0423 · $0.02/1k
Frequently asked
Which DeepSeek model is cheapest per run on these tasks?
DeepSeek V4 Flash 0423 at $0.10 per 1,000 runs, measured from real token usage, not list price.
How is the composite score calculated?
Each model gets a 0-1 score on cost, quality (mean judge score), speed, checks passed and brevity, relative to the other models in the family. Your slider weights combine them into a 0-100 composite. Cost, speed and brevity are log-scaled because prices span decades.
Are these the official prices?
They are the catalogue rates in models.json, synced hourly from OpenRouter's pricing API, multiplied by the tokens each run actually used. The $/1M columns show the rate; the cost columns show what the run cost.