OverpayingForAIPricing desk

Summarise · 3 checks · max 400 tokens · benched 2026-09-16

Muse Spark 1.3 vs Muse Glimmer 30B vs Llama 4 Maverick vs Llama 4 Scout on meeting notes to exactly three bullets

Meta models side by side on "Meeting notes to exactly three bullets": Llama 4 Scout scores 10/10; Llama 4 Scout is the cheapest answer scoring 8+ at $0.03 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a concise assistant. Output only what is asked for.

User

Turn these meeting notes into exactly three bullet points. Each bullet starts with "- " and is at most 20 words. Output nothing except the three bullets.

Notes: Launch of the new billing page moves from 18 Sept to 25 Sept because the tax calculation for EU customers still fails for VAT-exempt businesses. Priya owns the fix and will demo it Thursday. Marketing will hold the announcement email until the fix is verified. Separately, the team agreed to drop support for the legacy CSV export in Q1 next year and to notify affected customers in November.

Rubric for the judge: Three bullets that capture the delayed launch, the VAT fix owner, and the CSV export decision; nothing extra.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Muse Spark 1.3

meta/muse-spark-1.3

unjudged
failed: OpenRouter HTTP 403: This model requires you to complete the following before use: 18+ age confirmation. Confirm at https://openrouter.ai/settings/preferences.
Latency
0ms
Cost
$0
Per 1,000
$0

0 in · 0 out · 0 words · checks 0/3

  • regex
  • max_words
  • contains
(empty output)

Judge:

Muse Glimmer 30B

meta/muse-glimmer-30b

10/10
Latency
376ms
Cost
$0.00044
Per 1,000
$0.44

153 in · 358 out (312 reasoning) · 44 words · checks 3/3

  • regex
  • max_words
  • contains
- Billing page launch delayed from 18 Sept to 25 Sept due to EU VAT-exempt tax bug.
- Priya owns fix and will demo Thursday; marketing holds announcement email until verified.
- Legacy CSV export support drops Q1 next year; customers notified in November.

Judge: Output contains exactly three bullets, each under 20 words, captures all three required points (delayed launch, VAT fix owner, CSV export decision), and includes nothing extra.

Llama 4 Maverick

meta-llama/llama-4-maverick

10/10
Latency
285ms
Cost
$6.5e-5
Per 1,000
$0.06

141 in · 53 out · 45 words · checks 3/3

  • regex
  • max_words
  • contains
- Launch of new billing page delayed to 25 Sept due to EU tax calculation issues.
- Priya to demo tax fix on Thursday before marketing sends announcement email.
- Legacy CSV export support to be dropped in Q1 next year, customers notified in November.

Judge: Three bullets, each under 20 words, capturing all three required points (delayed launch, VAT fix owner/demo, CSV export decision) with no extra commentary.

Llama 4 Scout

meta-llama/llama-4-scout

10/10
top judge scorecheapest scoring 8+
Latency
7.9s
Cost
$2.9e-5
Per 1,000
$0.03

141 in · 51 out · 43 words · checks 3/3

  • regex
  • max_words
  • contains
- New billing page launch delayed to 25 Sept due to tax calculation issues for EU customers.
- Priya to demo tax calculation fix on Thursday for verification.
- Legacy CSV export support to be dropped in Q1, with customer notifications in November.

Judge: Three bullets, each under 20 words, capturing all three required points (delayed launch, VAT fix owner/demo, CSV export decision) with no extra commentary.

Frequently asked

What does this prompt test?

Summarise: Three bullets that capture the delayed launch, the VAT fix owner, and the CSV export decision; nothing extra. The deterministic checks are regex, max_words, contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Llama 4 Scout at $0.03 per 1,000 runs. If you need the top score, Llama 4 Scout at $0.03 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.