OverpayingForAIPricing desk

SQL · 6 checks · max 500 tokens · benched 2026-09-16

Qwen3.8 Max (0902) vs Qwen3.8 27B vs Qwen3.8 Flash vs Qwen3.7 Plus on top five customers by 2025 spend

Alibaba models side by side on "Top five customers by 2025 spend": Qwen3.8 Flash scores 10/10; Qwen3.8 Flash is the cheapest answer scoring 8+ at $0.45 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a database engineer. Reply with a single SQL statement and nothing else.

User

Tables: customers(id, name) and orders(id, customer_id, amount, ordered_at). Write one standard SQL query that returns each customer's name and their total order amount for calendar year 2025, highest total first, limited to the top 5. Use a JOIN, GROUP BY and ORDER BY.

Rubric for the judge: Correct join and aggregate, filters ordered_at to 2025, sorts descending, limits to 5, portable SQL.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Qwen3.8 Max (0902)

qwen/qwen3.8-max-0902

9/10
Latency
1.2s
Cost
$0.00147
Per 1,000
$1.47

122 in · 205 out (137 reasoning) · 32 words · checks 6/6

  • regex
  • regex
  • regex
  • regex
  • regex
  • contains
```sql
SELECT c.name, SUM(o.amount) AS total_amount
FROM customers c
JOIN orders o ON o.customer_id = c.id
WHERE EXTRACT(YEAR FROM o.ordered_at) = 2025
GROUP BY c.name
ORDER BY total_amount DESC
LIMIT 5;
```

Judge: Query correctly uses JOIN, GROUP BY, ORDER BY, and LIMIT with proper 2025 filtering; minor deduction because GROUP BY should include c.id to be fully standard-compliant (though c.name alone works in most systems).

Qwen3.8 27B

qwen/qwen3.8-27b

9/10
Latency
561ms
Cost
$0.00077
Per 1,000
$0.77

122 in · 292 out (174 reasoning) · 34 words · checks 6/6

  • regex
  • regex
  • regex
  • regex
  • regex
  • contains


```sql
SELECT c.name, SUM(o.amount) AS total_amount
FROM customers c
JOIN orders o ON c.id = o.customer_id
WHERE o.ordered_at >= '2025-01-01' AND o.ordered_at < '2026-01-01'
GROUP BY c.name
ORDER BY total_amount DESC
LIMIT 5;
```

Judge: Query correctly uses JOIN, GROUP BY, ORDER BY, and LIMIT with proper 2025 date filtering; minor deduction because GROUP BY should reference c.id instead of c.name to handle customers with identical names correctly, though the query will work in most SQL dialects.

Qwen3.8 Flash

qwen/qwen3.8-flash

10/10
top judge scorecheapest scoring 8+
Latency
1.7s
Cost
$0.00045
Per 1,000
$0.45

134 in · 921 out (847 reasoning) · 36 words · checks 5/6

  • regex
  • regex
  • regex
  • regex
  • regex
  • contains
SELECT c.name, SUM(o.amount) AS total_order_amount
FROM customers AS c
JOIN orders AS o ON c.id = o.customer_id
WHERE EXTRACT(YEAR FROM o.ordered_at) = 2025
GROUP BY c.id, c.name
ORDER BY total_order_amount DESC
FETCH FIRST 5 ROWS ONLY;

Judge: Query correctly joins customers and orders, filters 2025 using EXTRACT(YEAR FROM ...), groups by customer with proper aggregation, orders descending by total, limits to 5 rows using standard SQL FETCH syntax, and is fully portable across major databases.

Qwen3.7 Plus

qwen/qwen3.7-plus

1/10
failed: empty output (finish_reason length)
Latency
1.9s
Cost
$0.00221
Per 1,000
$2.21

96 in · 1700 out (1700 reasoning) · 0 words · checks 0/6

  • regex
  • regex
  • regex
  • regex
  • regex
  • contains
(empty output)

Judge: Output is completely empty; no SQL query was provided despite clear task requirements for a specific query with JOIN, GROUP BY, ORDER BY, and filtering.

Frequently asked

What does this prompt test?

SQL: Correct join and aggregate, filters ordered_at to 2025, sorts descending, limits to 5, portable SQL. The deterministic checks are regex, regex, regex, regex, regex, contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Qwen3.8 Flash at $0.45 per 1,000 runs. If you need the top score, Qwen3.8 Flash at $0.45 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.