GPT-5.5 vs GPT-5.4 vs GPT-5.4 mini vs GPT-5.4 Nano vs GPT-5.3-Codex on invoice text to json
OpenAI models side by side on "Invoice text to JSON": GPT-5.4 Nano scores 10/10; GPT-5.4 Nano is the cheapest answer scoring 8+ at $0.10 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
The prompt every model received
System
You are a data extraction engine. Reply with JSON only, no prose, no code fences.
User
Extract the fields below from the invoice text and return one JSON object with exactly these keys: vendor (string), invoice_number (string), total (number), currency (ISO 4217 string), due_date (YYYY-MM-DD string).
Invoice text:
Northwind Print Co.
Invoice INV-2041
Issued: 15 September 2026
Payment due: 15 October 2026
Items: 500 x A5 flyers @ 1.20 = 600.00; 4 x roll-up banners @ 171.13 = 684.50
Total due: USD 1,284.50
Thank you for your business.
Rubric for the judge: Valid JSON with the five keys, numeric total 1284.5, ISO due date, no extra keys or commentary.
Side by side
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
GPT-5.5
openai/gpt-5.5
10/10
Latency
334ms
Cost
$0.00306
Per 1,000
$3.06
162 in · 75 out (32 reasoning) · 3 words · checks 6/6
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.5), proper ISO 4217 currency code (USD), correctly formatted due date (2026-10-15), no extra keys or commentary.
GPT-5.4
openai/gpt-5.4
10/10
Latency
355ms
Cost
$0.00168
Per 1,000
$1.68
162 in · 85 out (42 reasoning) · 3 words · checks 6/6
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.5), proper ISO 4217 currency code (USD), correctly formatted due date (2026-10-15), no extra keys or commentary.
GPT-5.4 mini
openai/gpt-5.4-mini
10/10
Latency
314ms
Cost
$0.00049
Per 1,000
$0.49
162 in · 83 out (40 reasoning) · 3 words · checks 6/6
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.50), proper ISO 4217 currency code (USD), correctly formatted due date (2026-10-15), no extra keys or commentary.
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.5), proper ISO 4217 currency code (USD), correct ISO 8601 due date format (2026-10-15), no extra keys or commentary.
GPT-5.3-Codex
openai/gpt-5.3-codex
10/10
Latency
459ms
Cost
$0.00164
Per 1,000
$1.64
162 in · 97 out (54 reasoning) · 3 words · checks 6/6
Judge: Output is valid JSON with exactly the five required keys, correct numeric total (1284.5), proper ISO 4217 currency code (USD), correctly formatted due date (2026-10-15), no extra keys or commentary.
Frequently asked
▸What does this prompt test?
Extract JSON: Valid JSON with the five keys, numeric total 1284.5, ISO due date, no extra keys or commentary. The deterministic checks are json_valid, json_keys, contains, contains, contains, contains.
▸Which model should I pick for this task?
If the judge's bar of 8/10 is good enough for you, GPT-5.4 Nano at $0.10 per 1,000 runs. If you need the top score, GPT-5.4 Nano at $0.10 per 1,000 runs.
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.