Grok 4.6
x-ai/grok-4.6
- Latency
- 554ms
- Cost
- $0.00159
- Per 1,000
- $1.59
298 in · 198 out (197 reasoning) · 1 words · checks 1/1
- exact
90
Judge: The output is exactly '90' with no extraneous commentary, matching the rubric requirement perfectly.
Reasoning · 1 checks · max 800 tokens · benched 2026-09-16
xAI models side by side on "When does the faster train catch up?": Grok 4.6 scores 10/10; Grok 4.6 is the cheapest answer scoring 8+ at $1.59 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
System
Think carefully, then give only the final answer in the requested format.
User
Train A leaves a station at 09:10 travelling at a constant 80 km/h. Train B leaves the same station on the same track at 09:40 travelling at a constant 120 km/h in the same direction. How many minutes after 09:10 does train B catch up with train A? Reply with only the whole number of minutes.
Rubric for the judge: The single correct answer 90, with no working shown in the final output.
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
x-ai/grok-4.6
298 in · 198 out (197 reasoning) · 1 words · checks 1/1
90
Judge: The output is exactly '90' with no extraneous commentary, matching the rubric requirement perfectly.
x-ai/grok-4.5
298 in · 220 out (219 reasoning) · 1 words · checks 1/1
90
Judge: The output is exactly '90' with no extraneous commentary, matching the rubric requirement perfectly.
x-ai/grok-4.20
276 in · 715 out (558 reasoning) · 95 words · checks 0/1
90 The head start for Train A is 30 minutes or 0.5 hours, covering 80 * 0.5 = 40 km. The relative speed is 120 - 80 = 40 km/h, so catch-up time after B starts is 40 / 40 = 1 hour. Adding this to the 0.5-hour head start gives t = 1.5 hours or 90 minutes after 09:10. Alternatively, set up the equation with t in hours after 09:10: 80t = 120(t - 0.5). Solving yields 80t = 120t - 60, so 60 = 40t and t = 1.5 hours, confirming 90 minutes.
Judge: The answer '90' is correct, but the output includes extensive working and explanation when the rubric explicitly requires only the whole number with no working shown.
x-ai/grok-4.3
282 in · 441 out (273 reasoning) · 108 words · checks 0/1
90 The head start of Train A is 30 minutes, or 0.5 h, during which it covers \(80 \times 0.5 = 40\) km. The relative speed advantage of Train B is \(120 - 80 = 40\) km/h, so the time required to close the 40 km gap is \(40 / 40 = 1\) h. Adding the 30-minute head start to this 1 h interval yields a total of 90 minutes after 09:10. Verification: in 90 min (1.5 h) Train A travels \(80 \times 1.5 = 120\) km; in the remaining 60 min Train B travels \(120 \times 1 = 120\) km, confirming they meet at the same point.
Judge: The answer '90' is correct, but the rubric explicitly requires 'no working shown in the final output' and the model provided extensive working and verification, violating the format requirement.
Reasoning: The single correct answer 90, with no working shown in the final output. The deterministic checks are exact.
If the judge's bar of 8/10 is good enough for you, Grok 4.6 at $1.59 per 1,000 runs. If you need the top score, Grok 4.6 at $1.59 per 1,000 runs.
Run it yourself
Re-run this prompt on xAI with your own OpenRouter key, or tweak the wording and see what changes.
More xAI prompts
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.
AI cost intelligence
Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.
Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.