Muse Spark 1.3
meta/muse-spark-1.3
- Latency
- 0ms
- Cost
- $0
- Per 1,000
- $0
0 in · 0 out · 0 words · checks 0/7
- json_valid
- json_keys
- regex
- contains
- regex
- contains
- contains
(empty output)
Judge: —
Tool call · 7 checks · max 400 tokens · benched 2026-09-16
Meta models side by side on "Emit a calendar tool call": Llama 4 Scout scores 10/10; Llama 4 Scout is the cheapest answer scoring 8+ at $0.03 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
System
You are a function-calling engine. Reply with JSON only, no prose, no code fences.
User
You have one tool:
{"name": "create_event", "parameters": {"title": "string", "start": "ISO 8601 date-time without timezone, e.g. 2026-01-31T09:00", "duration_minutes": "integer", "attendees": "array of email strings"}}
Produce the tool call for this user message as a JSON object of the shape {"name": string, "arguments": object}.
User: Book a 45-minute design review with [email protected] and [email protected] on 2 October 2026 at 3pm.Rubric for the judge: Correct tool name, ISO start 2026-10-02T15:00, integer duration 45, both attendees in an array, a sensible title.
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
meta/muse-spark-1.3
0 in · 0 out · 0 words · checks 0/7
(empty output)
Judge: —
meta/muse-glimmer-30b
168 in · 165 out (91 reasoning) · 13 words · checks 7/7
{"name": "create_event", "arguments": {"title": "design review", "start": "2026-10-02T15:00", "duration_minutes": 45, "attendees": ["[email protected]", "[email protected]"]}}Judge: All rubric criteria met: correct tool name, proper ISO 8601 start datetime, integer duration of 45 minutes, both attendees in array format, and sensible title.
meta-llama/llama-4-maverick
156 in · 55 out · 13 words · checks 7/7
{"name": "create_event", "arguments": {"title": "Design Review", "start": "2026-10-02T15:00", "duration_minutes": 45, "attendees": ["[email protected]", "[email protected]"]}}Judge: All rubric criteria met: correct tool name, proper ISO 8601 start datetime, integer duration of 45 minutes, both attendees in array format, and sensible title.
meta-llama/llama-4-scout
156 in · 55 out · 13 words · checks 7/7
{"name": "create_event", "arguments": {"title": "design review", "start": "2026-10-02T15:00", "duration_minutes": 45, "attendees": ["[email protected]", "[email protected]"]}}Judge: All rubric criteria met: correct tool name, proper ISO 8601 start datetime, integer duration of 45 minutes, both attendees in array format, and sensible title.
Tool call: Correct tool name, ISO start 2026-10-02T15:00, integer duration 45, both attendees in an array, a sensible title. The deterministic checks are json_valid, json_keys, regex, contains, regex, contains, contains.
If the judge's bar of 8/10 is good enough for you, Llama 4 Scout at $0.03 per 1,000 runs. If you need the top score, Llama 4 Scout at $0.03 per 1,000 runs.
Run it yourself
Re-run this prompt on Meta with your own OpenRouter key, or tweak the wording and see what changes.
More Meta prompts
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.
AI cost intelligence
Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.
Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.