OverpayingForAIPricing desk

Reasoning · 1 checks · max 800 tokens · benched 2026-09-16

Grok 4.6 vs Grok 4.5 vs Grok 4.20 vs Grok 4.3 on count sellable units

xAI models side by side on "Count sellable units": Grok 4.5 scores 10/10; Grok 4.5 is the cheapest answer scoring 8+ at $1.59 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

Think carefully, then give only the final answer in the requested format.

User

A warehouse receives 3 pallets. Each pallet holds 48 boxes and each box holds 12 units. On inspection, 15% of the boxes are damaged; round the number of damaged boxes down to a whole box. The damaged boxes are discarded. Then 7 undamaged boxes are sent to the quality lab and are not sold. How many units are left to sell? Reply with only the number.

Rubric for the judge: The single correct answer 1392 (144 boxes, 21 damaged, 7 to the lab, 116 x 12).

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

Grok 4.6

x-ai/grok-4.6

10/10
Latency
573ms
Cost
$0.00187
Per 1,000
$1.87

307 in · 209 out (207 reasoning) · 1 words · checks 1/1

  • exact
1392

Judge: The output matches the correct answer of 1392 units exactly, provided in the required format of a bare number only.

Grok 4.5

x-ai/grok-4.5

10/10
top judge scorecheapest scoring 8+
Latency
293ms
Cost
$0.00159
Per 1,000
$1.59

307 in · 199 out (197 reasoning) · 1 words · checks 1/1

  • exact
1392

Judge: The output matches the correct answer of 1392 units exactly, provided in the required format of a bare number only.

Grok 4.20

x-ai/grok-4.20

3/10
Latency
392ms
Cost
$0.00190
Per 1,000
$1.90

285 in · 673 out (525 reasoning) · 111 words · checks 0/1

  • exact
1392


The total number of boxes is calculated as 3 pallets × 48 boxes per pallet, resulting in 144 boxes. Fifteen percent of 144 boxes equals 21.6, which is rounded down to 21 damaged boxes using the floor function. These 21 damaged boxes are discarded, leaving 144 − 21 = 123 undamaged boxes. After sending 7 undamaged boxes to the quality lab, 123 − 7 = 116 boxes remain available for sale. Each box contains 12 units, so the total units left to sell are 116 × 12 = 1392. All steps follow the problem constraints exactly, with no intermediate rounding or assumptions beyond the specified floor operation on damaged boxes.

Judge: Output correctly provides the answer '1392' but violates the explicit instruction to 'Reply with only the number' by including extensive explanatory commentary.

Grok 4.3

x-ai/grok-4.3

3/10
Latency
451ms
Cost
$0.00117
Per 1,000
$1.17

291 in · 378 out (189 reasoning) · 141 words · checks 0/1

  • exact
1392 

The calculation begins by determining the total number of boxes received: 3 pallets multiplied by 48 boxes per pallet yields 144 boxes. Next, 15% of these boxes are identified as damaged, which is 0.15 × 144 = 21.6; this value is rounded down to the nearest whole number, resulting in 21 damaged boxes. These damaged boxes are removed, leaving 144 − 21 = 123 undamaged boxes. From the remaining undamaged boxes, 7 are allocated to the quality lab and therefore excluded, producing 123 − 7 = 116 boxes available for sale. Finally, each of these boxes contains 12 units, so the total number of sellable units is obtained by multiplying 116 × 12 = 1392. This sequence follows the problem constraints exactly, applying the floor operation only to the damaged-box count and subtracting the specified quantities in the stated order.

Judge: The output provides the correct numerical answer (1392) but violates the explicit instruction to 'Reply with only the number' by including extensive explanatory commentary.

Frequently asked

What does this prompt test?

Reasoning: The single correct answer 1392 (144 boxes, 21 damaged, 7 to the lab, 116 x 12). The deterministic checks are exact.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, Grok 4.5 at $1.59 per 1,000 runs. If you need the top score, Grok 4.5 at $1.59 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.