Muse Spark 1.3
meta/muse-spark-1.3
- Latency
- 0ms
- Cost
- $0
- Per 1,000
- $0
0 in · 0 out · 0 words · checks 0/4
- contains
- regex
- not_contains
- not_contains
(empty output)
Judge: —
Code fix · 4 checks · max 600 tokens · benched 2026-09-16
Meta models side by side on "Fix an off-by-one in Python": Llama 4 Scout scores 10/10; Llama 4 Scout is the cheapest answer scoring 8+ at $0.02 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
System
You are a senior engineer. Reply with code only unless asked otherwise.
User
This function should return the arithmetic mean of a non-empty list of numbers, but its result is wrong. Fix the bug with the smallest possible change and return the complete corrected function only, in a Python code block.
def average(nums):
total = 0
for n in nums:
total += n
return total / (len(nums) - 1)Rubric for the judge: Returns the full function with the divisor corrected to len(nums) and no other behavioural changes.
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
meta/muse-spark-1.3
0 in · 0 out · 0 words · checks 0/4
(empty output)
Judge: —
meta/muse-glimmer-30b
117 in · 129 out (107 reasoning) · 18 words · checks 4/4
```python
def average(nums):
total = 0
for n in nums:
total += n
return total / len(nums)
```Judge: The output correctly fixes the bug by changing the divisor from len(nums) - 1 to len(nums), returns the complete function in a Python code block, and makes no other changes.
meta-llama/llama-4-maverick
105 in · 33 out · 18 words · checks 4/4
```python
def average(nums):
total = 0
for n in nums:
total += n
return total / len(nums)
```Judge: The output correctly fixes the bug by changing the divisor from len(nums) - 1 to len(nums), returns the complete function in a Python code block, and makes no other changes.
meta-llama/llama-4-scout
105 in · 33 out · 18 words · checks 4/4
```python
def average(nums):
total = 0
for n in nums:
total += n
return total / len(nums)
```Judge: The output correctly fixes the bug by changing the divisor from len(nums) - 1 to len(nums), returns the complete function in a Python code block, and makes no other changes.
Code fix: Returns the full function with the divisor corrected to len(nums) and no other behavioural changes. The deterministic checks are contains, regex, not_contains, not_contains.
If the judge's bar of 8/10 is good enough for you, Llama 4 Scout at $0.02 per 1,000 runs. If you need the top score, Llama 4 Scout at $0.02 per 1,000 runs.
Run it yourself
Re-run this prompt on Meta with your own OpenRouter key, or tweak the wording and see what changes.
More Meta prompts
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.
AI cost intelligence
Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.
Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.