Grok 4.6 vs Grok 4.5 vs Grok 4.20 vs Grok 4.3 on invoice text to json
xAI models side by side on "Invoice text to JSON": Grok 4.3 scores 10/10; Grok 4.3 is the cheapest answer scoring 8+ at $1.29 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
The prompt every model received
System
You are a data extraction engine. Reply with JSON only, no prose, no code fences.
User
Extract the fields below from the invoice text and return one JSON object with exactly these keys: vendor (string), invoice_number (string), total (number), currency (ISO 4217 string), due_date (YYYY-MM-DD string).
Invoice text:
Northwind Print Co.
Invoice INV-2041
Issued: 15 September 2026
Payment due: 15 October 2026
Items: 500 x A5 flyers @ 1.20 = 600.00; 4 x roll-up banners @ 171.13 = 684.50
Total due: USD 1,284.50
Thank you for your business.
Rubric for the judge: Valid JSON with the five keys, numeric total 1284.5, ISO due date, no extra keys or commentary.
Side by side
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
Grok 4.6
x-ai/grok-4.6
10/10
Latency
457ms
Cost
$0.00171
Per 1,000
$1.71
368 in · 163 out (119 reasoning) · 3 words · checks 6/6
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.5), proper ISO 4217 currency code (USD), correctly formatted due date (2026-10-15), no extra keys or commentary.
Grok 4.5
x-ai/grok-4.5
10/10
Latency
490ms
Cost
$0.00138
Per 1,000
$1.38
368 in · 144 out (100 reasoning) · 3 words · checks 6/6
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.5), proper ISO 4217 currency code (USD), correctly formatted due date (2026-10-15), no extra keys or commentary.
Grok 4.20
x-ai/grok-4.20
10/10
Latency
255ms
Cost
$0.00257
Per 1,000
$2.57
341 in · 910 out (857 reasoning) · 14 words · checks 6/6
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.5), proper ISO 4217 currency code (USD), correct ISO 8601 due date format (2026-10-15), no extra keys or commentary.
Grok 4.3
x-ai/grok-4.3
10/10
top judge scorecheapest scoring 8+
Latency
257ms
Cost
$0.00129
Per 1,000
$1.29
347 in · 422 out (369 reasoning) · 14 words · checks 6/6
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.5), proper ISO 4217 currency code (USD), correct ISO 8601 due date format (2026-10-15), no extra keys or commentary.
Frequently asked
▸What does this prompt test?
Extract JSON: Valid JSON with the five keys, numeric total 1284.5, ISO due date, no extra keys or commentary. The deterministic checks are json_valid, json_keys, contains, contains, contains, contains.
▸Which model should I pick for this task?
If the judge's bar of 8/10 is good enough for you, Grok 4.3 at $1.29 per 1,000 runs. If you need the top score, Grok 4.3 at $1.29 per 1,000 runs.
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.