Muse Spark 1.3
meta/muse-spark-1.3
- Latency
- 0ms
- Cost
- $0
- Per 1,000
- $0
0 in · 0 out · 0 words · checks 0/5
- regex
- regex
- regex
- regex
- max_words
(empty output)
Judge: —
Classify · 5 checks · max 400 tokens · benched 2026-09-16
Meta models side by side on "Sentiment of five reviews": Llama 4 Scout scores 10/10; Llama 4 Scout is the cheapest answer scoring 8+ at $0.02 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
System
You are a classifier. Output exactly the requested lines and nothing else.
User
Label each review as positive, neutral or negative. Output five lines in the form "<number>: <label>" and nothing else. 1. Battery lasts two full days and the screen is gorgeous. Would buy again. 2. It arrived on Tuesday in a brown box. 3. Stopped charging after three weeks and support never replied. 4. Does what it says. Nothing special, nothing wrong. 5. The strap broke on day one and the app crashes constantly.
Rubric for the judge: Five correctly labelled lines in the exact format; review 2 and 4 are neutral; no explanation.
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
meta/muse-spark-1.3
0 in · 0 out · 0 words · checks 0/5
(empty output)
Judge: —
meta/muse-glimmer-30b
135 in · 132 out (116 reasoning) · 10 words · checks 5/5
1: positive 2: neutral 3: negative 4: neutral 5: negative
Judge: All five reviews correctly labeled in exact format with no extra commentary; reviews 2 and 4 properly identified as neutral.
meta-llama/llama-4-maverick
123 in · 20 out · 10 words · checks 5/5
1: positive 2: neutral 3: negative 4: neutral 5: negative
Judge: All five reviews correctly labeled in exact format with no extra commentary; reviews 2 and 4 properly identified as neutral.
meta-llama/llama-4-scout
123 in · 20 out · 10 words · checks 5/5
1: positive 2: neutral 3: negative 4: neutral 5: negative
Judge: All five reviews correctly labeled in exact format with no extra commentary; reviews 2 and 4 properly identified as neutral.
Classify: Five correctly labelled lines in the exact format; review 2 and 4 are neutral; no explanation. The deterministic checks are regex, regex, regex, regex, max_words.
If the judge's bar of 8/10 is good enough for you, Llama 4 Scout at $0.02 per 1,000 runs. If you need the top score, Llama 4 Scout at $0.02 per 1,000 runs.
Run it yourself
Re-run this prompt on Meta with your own OpenRouter key, or tweak the wording and see what changes.
More Meta prompts
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.
AI cost intelligence
Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.
Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.