OverpayingForAIPricing desk

Lesson 2 of 5 · 7 min read · Beginner

Reading a leaderboard: cost, checks, judge, latency

Walk through one real family leaderboard, Anthropic on the bench of 2026-09-12, column by column. Cost per 1,000 runs is the number that matters, and the cheapest row is closer to the top than you would guess.

In this lesson you will

  • Read each column of a real leaderboard and know which to trust first
  • Convert cost per run into cost per 1,000 runs and a monthly figure
  • Spot the trade-off the cheapest row is making

Here is the Anthropic family from the bench of 2026-09-12: 4 models, 24 prompts, 96 runs, judged by Claude Haiku 4.5, total spend $0.19. Every number below is copied from the results, not rounded to make a point.

Anthropic family, bench of 2026-09-12 (prompts v2026-09-12.1, judge anthropic/claude-haiku-4.5, reasoning effort low). Rates in $ per 1M tokens; cost per 1,000 runs is the mean real cost of one run × 1,000.
Model$ in / out per 1MCost per 1,000 runsChecks passedJudge (0–10)Mean latency
Claude Sonnet 5$2 / $10$1.0396.0%9.791,183 ms
Claude Sonnet 4.6$3 / $15$1.2097.0%9.50864 ms
Claude Opus 5$5 / $25$2.5998.0%9.461,420 ms
Claude Haiku 4.5$1 / $5$0.5292.9%9.46746 ms

Column by column

  • Cost per 1,000 runs — read this first. Haiku 4.5 at $0.52 is a fifth of Opus 5 at $2.59 for the same 24 tasks. Multiply by your monthly volume: at 10,000 runs a month that is $5.20 against $25.90.
  • Checks passed — the hard rules. Opus leads at 98.0%; Haiku is lowest at 92.9%. That 5-point gap is the honest cost of going cheap, and lesson 4 shows what it looks like in practice.
  • Judge — Sonnet 5 scores highest at 9.79 while costing less than Opus. Judge scores cluster near the top for good models; a difference of 0.3 is noise, a difference of 3 is a real failure somewhere.
  • Latency — Haiku answers in 746 ms, Opus in 1,420 ms. For one question nobody cares. For a loop of a thousand, that is 11 minutes versus 24.

Knowledge check

From the table, roughly how much would 20,000 runs a month cost on Claude Haiku 4.5 versus Claude Opus 5?

Lesson FAQ

What does cost per 1,000 runs mean?

The average real cost of one run of the 24 bench prompts, multiplied by 1,000. It is computed from the tokens each model actually used at its catalogue rate, so it already reflects that some models write longer answers than others.

Is a higher judge score always better?

Within a family, mostly. Across families, the judge model has a house style and can favour answers that sound like its own. Read checks passed first; they are pass/fail rules with no taste involved.

Why is Claude Sonnet 4.6 more expensive than Sonnet 5?

Sonnet 4.6 lists at $3 / $15 and Sonnet 5 at $2 / $10 in the catalogue on 2026-09-12; newer models are often priced below the ones they replace. This is one reason to re-read the leaderboard when a new model lands.

Which Claude model is cheapest that still works?

On the bench of 2026-09-12, Claude Haiku 4.5 passed 92.9% of checks at $0.52 per 1,000 runs. Whether that is 'works' depends on your task; lesson 4 shows exactly which prompts it missed.

Finished reading?

Mark it done to track your progress through the course.

Save progress across devices

Get a private link that restores your lessons on any device. Email is optional and only used to send you the link.

Compare, calculate, decide

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.