Muse Spark 1.3
meta/muse-spark-1.3
- Latency
- 0ms
- Cost
- $0
- Per 1,000
- $0
0 in · 0 out · 0 words · checks 0/6
- json_valid
- regex
- regex
- regex
- regex
- regex
(empty output)
Judge: —
Classify · 6 checks · max 400 tokens · benched 2026-09-16
Meta models side by side on "Label six support tickets": Llama 4 Scout scores 10/10; Llama 4 Scout is the cheapest answer scoring 8+ at $0.03 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
System
You are a classifier. Use only the allowed labels. Reply with JSON only, no prose, no code fences.
User
Classify each ticket into exactly one label from this set: billing, bug, feature_request, account_access. Return a JSON object mapping ticket id to label. t1: "I was invoiced for 12 seats but we only use 8." t2: "The export button does nothing when the report has more than 1,000 rows." t3: "Every time I open the dashboard on Safari the charts render blank." t4: "It would be great if we could schedule reports to email weekly." t5: "I reset my password twice and still cannot log in; it says my account is locked." t6: "Can you add a dark mode to the mobile app?"
Rubric for the judge: All six tickets get the right label from the fixed set, returned as a compact JSON object with no commentary.
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
meta/muse-spark-1.3
0 in · 0 out · 0 words · checks 0/6
(empty output)
Judge: —
meta/muse-glimmer-30b
192 in · 183 out (143 reasoning) · 12 words · checks 6/6
{"t1": "billing", "t2": "bug", "t3": "bug", "t4": "feature_request", "t5": "account_access", "t6": "feature_request"}Judge: All six tickets are correctly classified into the specified label set with proper JSON formatting and no extraneous commentary.
meta-llama/llama-4-maverick
180 in · 46 out · 12 words · checks 6/6
{"t1": "billing", "t2": "bug", "t3": "bug", "t4": "feature_request", "t5": "account_access", "t6": "feature_request"}Judge: All six tickets are correctly classified into the specified label set with proper JSON formatting and no extraneous commentary.
meta-llama/llama-4-scout
180 in · 56 out · 16 words · checks 6/6
```
{
"t1": "billing",
"t2": "bug",
"t3": "bug",
"t4": "feature_request",
"t5": "account_access",
"t6": "feature_request"
}
```Judge: All six tickets are correctly classified into the appropriate labels from the fixed set, returned as a compact JSON object with no extra commentary.
Classify: All six tickets get the right label from the fixed set, returned as a compact JSON object with no commentary. The deterministic checks are json_valid, regex, regex, regex, regex, regex.
If the judge's bar of 8/10 is good enough for you, Llama 4 Scout at $0.03 per 1,000 runs. If you need the top score, Llama 4 Scout at $0.03 per 1,000 runs.
Run it yourself
Re-run this prompt on Meta with your own OpenRouter key, or tweak the wording and see what changes.
More Meta prompts
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.
AI cost intelligence
Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.
Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.