OverpayingForAIPricing desk

Reasoning · 1 checks · max 800 tokens · benched 2026-09-16

Claude Opus 5 vs Claude Sonnet 5 vs Claude Sonnet 4.6 vs Claude Haiku 4.5 on when does the faster train catch up?

Anthropic models side by side on "When does the faster train catch up?": Claude Sonnet 4.6 scores 10/10; Claude Sonnet 4.6 is the cheapest answer scoring 8+ at $0.38 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

Think carefully, then give only the final answer in the requested format.

User

Train A leaves a station at 09:10 travelling at a constant 80 km/h. Train B leaves the same station on the same track at 09:40 travelling at a constant 120 km/h in the same direction. How many minutes after 09:10 does train B catch up with train A? Reply with only the whole number of minutes.

Rubric for the judge: The single correct answer 90, with no working shown in the final output.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Claude Opus 5

anthropic/claude-opus-5

10/10
Latency
1.2s
Cost
$0.00211
Per 1,000
$2.11

127 in · 59 out (56 reasoning) · 1 words · checks 1/1

  • exact
90

Judge: The output is exactly '90' with no extraneous commentary, matching the rubric requirement perfectly.

Claude Sonnet 5

anthropic/claude-sonnet-5

10/10
Latency
1.4s
Cost
$0.00094
Per 1,000
$0.94

127 in · 69 out (66 reasoning) · 1 words · checks 1/1

  • exact
90

Judge: The output is exactly '90' with no extraneous commentary, matching the rubric requirement perfectly.

Claude Sonnet 4.6

anthropic/claude-sonnet-4.6

10/10
top judge scorecheapest scoring 8+
Latency
886ms
Cost
$0.00038
Per 1,000
$0.38

102 in · 5 out · 1 words · checks 1/1

  • exact
90

Judge: The output is exactly '90' with no extraneous commentary, matching the rubric requirement perfectly.

Claude Haiku 4.5

anthropic/claude-haiku-4.5

3/10
Latency
711ms
Cost
$0.00165
Per 1,000
$1.65

101 in · 310 out · 166 words · checks 0/1

  • exact
I need to find when Train B catches up with Train A.

**Setup:**
- Train A: leaves at 09:10, speed = 80 km/h
- Train B: leaves at 09:40, speed = 120 km/h
- Train B leaves 30 minutes after Train A

**Distance of Train A when Train B starts:**
In 30 minutes, Train A travels: 80 km/h × 0.5 h = 40 km

**Relative speed:**
Train B is catching up at: 120 - 80 = 40 km/h

**Time for Train B to catch up:**
Time = Distance / Relative speed = 40 km / 40 km/h = 1 hour = 60 minutes

**Time after 09:10:**
Train B starts 30 minutes after 09:10 and takes 60 minutes to catch up.
Total time after 09:10 = 30 + 60 = 90 minutes

**Verification:**
- Train A position at 90 minutes: 80 × (90/60) = 80 × 1.5 = 120 km
- Train B position at 60 minutes after 09:40: 120 × (60/60) = 120 km ✓

90

Judge: The answer of 90 is correct, but the rubric explicitly requires 'only the whole number of minutes' with 'no working shown in the final output', and the model provided extensive working before the answer.

Frequently asked

What does this prompt test?

Reasoning: The single correct answer 90, with no working shown in the final output. The deterministic checks are exact.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Claude Sonnet 4.6 at $0.38 per 1,000 runs. If you need the top score, Claude Sonnet 4.6 at $0.38 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.