Here is the Anthropic family from the bench of 2026-09-12: 4 models, 24 prompts, 96 runs, judged by Claude Haiku 4.5, total spend $0.19. Every number below is copied from the results, not rounded to make a point.
| Model | $ in / out per 1M | Cost per 1,000 runs | Checks passed | Judge (0–10) | Mean latency |
|---|---|---|---|---|---|
| Claude Sonnet 5 | $2 / $10 | $1.03 | 96.0% | 9.79 | 1,183 ms |
| Claude Sonnet 4.6 | $3 / $15 | $1.20 | 97.0% | 9.50 | 864 ms |
| Claude Opus 5 | $5 / $25 | $2.59 | 98.0% | 9.46 | 1,420 ms |
| Claude Haiku 4.5 | $1 / $5 | $0.52 | 92.9% | 9.46 | 746 ms |
Column by column
- Cost per 1,000 runs — read this first. Haiku 4.5 at $0.52 is a fifth of Opus 5 at $2.59 for the same 24 tasks. Multiply by your monthly volume: at 10,000 runs a month that is $5.20 against $25.90.
- Checks passed — the hard rules. Opus leads at 98.0%; Haiku is lowest at 92.9%. That 5-point gap is the honest cost of going cheap, and lesson 4 shows what it looks like in practice.
- Judge — Sonnet 5 scores highest at 9.79 while costing less than Opus. Judge scores cluster near the top for good models; a difference of 0.3 is noise, a difference of 3 is a real failure somewhere.
- Latency — Haiku answers in 746 ms, Opus in 1,420 ms. For one question nobody cares. For a loop of a thousand, that is 11 minutes versus 24.
Knowledge check
From the table, roughly how much would 20,000 runs a month cost on Claude Haiku 4.5 versus Claude Opus 5?