Muse Spark 1.3
meta/muse-spark-1.3
- Latency
- 0ms
- Cost
- $0
- Per 1,000
- $0
0 in · 0 out · 0 words · checks 0/3
- max_words
- contains
- not_contains
(empty output)
Judge: —
Summarise · 3 checks · max 400 tokens · benched 2026-09-16
Meta models side by side on "Summarise a support thread in 60 words": Llama 4 Scout scores 9/10; Llama 4 Scout is the cheapest answer scoring 8+ at $0.04 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
System
You are a concise assistant. Follow length limits exactly and output plain text only.
User
Summarise the following customer support thread in no more than 60 words. Mention what the customer wanted and how it was resolved. --- Customer (Mon 09:12): Hi, I was charged twice for my Team plan on 1 September. Order references are 88213 and 88214. I only have one workspace. Can you fix this? Agent (Mon 10:40): Sorry about that. I can see the duplicate charge. Could you confirm the last four digits of the card so I can raise the refund? Customer (Mon 11:02): 4471. Agent (Mon 11:30): Thanks. I have refunded order 88214 in full (USD 96.00). It should appear on your statement within 5 business days. I have also added a note so it cannot recur next cycle. Customer (Mon 11:35): Great, thanks for the quick turnaround. ---
Rubric for the judge: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits.
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
meta/muse-spark-1.3
0 in · 0 out · 0 words · checks 0/3
(empty output)
Judge: —
meta/muse-glimmer-30b
233 in · 273 out (186 reasoning) · 49 words · checks 2/3
Customer wanted a refund for a duplicate Team plan charge on 1 September, orders 88213 and 88214, with one workspace. Agent verified card ending 4471, refunded order 88214 in full USD 96.00 within 5 business days, and added a note to prevent recurrence. Customer thanked for the quick turnaround.
Judge: Output exceeds 60-word limit (62 words) and leaks sensitive card digits (4471), violating the rubric's explicit requirement to not leak card information.
meta-llama/llama-4-maverick
221 in · 58 out · 45 words · checks 3/3
The customer was charged twice for their Team plan. The issue was resolved when the agent refunded the duplicate charge (order 88214, USD 96.00) after confirming the customer's card details, and added a note to prevent it from happening again in the next billing cycle.
Judge: Accurate summary covering what the customer wanted (fix double charge) and resolution (refund + prevention measure), well under 60 words (42 words), no card digits leaked, minor deduction for not explicitly stating the customer had one workspace.
meta-llama/llama-4-scout
221 in · 62 out · 52 words · checks 3/3
A customer was charged twice for their Team plan. The agent identified the duplicate charge, refunded the incorrect charge of $96, and added a note to prevent future recurrence. The issue was resolved within an hour, and the customer was informed the refund would appear on their statement within 5 business days.
Judge: Accurate summary covering what the customer wanted (duplicate charge fixed) and resolution (refund issued, prevention note added), well under 60 words (52 words), no card digits leaked, though could have been slightly more explicit about the two order references.
Summarise: Accurate, complete summary of the double charge and its refund, within 60 words, no leaked card digits. The deterministic checks are max_words, contains, not_contains.
If the judge's bar of 8/10 is good enough for you, Llama 4 Scout at $0.04 per 1,000 runs. If you need the top score, Llama 4 Scout at $0.04 per 1,000 runs.
Run it yourself
Re-run this prompt on Meta with your own OpenRouter key, or tweak the wording and see what changes.
More Meta prompts
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.
AI cost intelligence
Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.
Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.