Vendor-neutral · Beginner · 5 lessons · about 33 minutes · updated 2026-09-12
Compare AI Models on the Same Prompt: Reading the Playground
One prompt, a whole model family, real costs. Five lessons on reading a leaderboard, running your own test, and knowing when the cheap model is enough.
For people who have heard that cheaper models are 'good enough' and want to see the evidence before believing it. No coding required; the only tools are the playground on this site and, optionally, an OpenRouter key.
Course outline
- 1Why one prompt across a family is the fair testModel comparisons online are mostly vibes. The playground runs one identical prompt through every model in a family and records what each did, how long it took and what it cost. Learn what that removes and what it cannot.5 min · 3 objectives
- 2Reading a leaderboard: cost, checks, judge, latencyWalk through one real family leaderboard, Anthropic on the bench of 2026-09-12, column by column. Cost per 1,000 runs is the number that matters, and the cheapest row is closer to the top than you would guess.7 min · 3 objectives
- 3Run your own prompt three ways: keyless, your key, ProThe bench prompts are not your prompt. Run yours: keyless on the site (free, cheapest models, three a day), with your own OpenRouter key (any model, your bill), or on Pro (the whole catalogue on our key).8 min · 3 objectives
- 4When the cheap model is enough, and when to pay for the flagshipTwo real results from the bench of 2026-09-12: a support-thread summary where the cheapest Claude beat the flagship at a seventh of the cost, and arithmetic prompts where the cheapest Mistral got every answer wrong. Learn to tell the two cases apart.7 min · 3 objectives
- 5Turn a result into a decisionA playground result is only useful once it becomes a sentence: for this task, this model, because of this evidence, at this cost. Write yours, price it for a month, and know when to re-test.6 min · 3 objectives
Frequently asked
What is the playground?
A page on this site that sends the same prompt to every model in a vendor's family and records the answer, whether it passed the prompt's checks, a judge score, the real cost per run and the latency. It covers 24 fixed prompts per family, plus a run page for your own prompt.
Can I use it without an account?
Yes. Reading every result is free, and keyless runs of your own prompt are free three times a day on the cheapest models. Pasting your own OpenRouter key unlocks any model at your own cost; Pro runs the whole catalogue on the site's key.
Are the numbers in this course real?
Yes. Every leaderboard figure, per-run cost and quoted answer comes from the bench of 2026-09-12 in the site's results files. They will drift as models are replaced; the playground page shows the current bench date.
Which model is cheapest that still works?
It depends on the task, which is the point of the course. On the bench of 2026-09-12 the cheapest Anthropic model handled rule-based summaries as well as the flagship, and the cheapest Mistral model failed every arithmetic prompt. Test your own task before believing either story.