OverpayingForAIPricing desk

Tool call · 7 checks · max 400 tokens · benched 2026-09-16

GPT-5.5 vs GPT-5.4 vs GPT-5.4 mini vs GPT-5.4 Nano vs GPT-5.3-Codex on emit a calendar tool call

OpenAI models side by side on "Emit a calendar tool call": GPT-5.4 Nano scores 10/10; GPT-5.4 Nano is the cheapest answer scoring 8+ at $0.14 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.

The prompt every model received

System

You are a function-calling engine. Reply with JSON only, no prose, no code fences.

User

You have one tool:
{"name": "create_event", "parameters": {"title": "string", "start": "ISO 8601 date-time without timezone, e.g. 2026-01-31T09:00", "duration_minutes": "integer", "attendees": "array of email strings"}}

Produce the tool call for this user message as a JSON object of the shape {"name": string, "arguments": object}.

User: Book a 45-minute design review with [email protected] and [email protected] on 2 October 2026 at 3pm.

Rubric for the judge: Correct tool name, ISO start 2026-10-02T15:00, integer duration 45, both attendees in an array, a sensible title.

Side by side

Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.

GPT-5.5

openai/gpt-5.5

10/10
Latency
432ms
Cost
$0.00316
Per 1,000
$3.16

152 in · 80 out (31 reasoning) · 2 words · checks 7/7

  • json_valid
  • json_keys
  • regex
  • contains
  • regex
  • contains
  • contains
{"name":"create_event","arguments":{"title":"Design review","start":"2026-10-02T15:00","duration_minutes":45,"attendees":["[email protected]","[email protected]"]}}

Judge: All parameters are correct: tool name is 'create_event', ISO start datetime is '2026-10-02T15:00', duration is 45 minutes as an integer, both attendees are in an array, and title 'Design review' is sensible and matches the user request.

GPT-5.4

openai/gpt-5.4

10/10
Latency
411ms
Cost
$0.00166
Per 1,000
$1.66

152 in · 85 out (36 reasoning) · 2 words · checks 7/7

  • json_valid
  • json_keys
  • regex
  • contains
  • regex
  • contains
  • contains
{"name":"create_event","arguments":{"title":"Design review","start":"2026-10-02T15:00","duration_minutes":45,"attendees":["[email protected]","[email protected]"]}}

Judge: All parameters are correct: tool name is 'create_event', ISO start datetime is '2026-10-02T15:00', duration is 45 minutes as an integer, both attendees are in an array, and title 'Design review' is sensible and matches the user request.

GPT-5.4 mini

openai/gpt-5.4-mini

10/10
Latency
411ms
Cost
$0.00048
Per 1,000
$0.48

152 in · 81 out (32 reasoning) · 2 words · checks 7/7

  • json_valid
  • json_keys
  • regex
  • contains
  • regex
  • contains
  • contains
{"name":"create_event","arguments":{"title":"Design review","start":"2026-10-02T15:00","duration_minutes":45,"attendees":["[email protected]","[email protected]"]}}

Judge: All parameters are correct: tool name is 'create_event', ISO start datetime is '2026-10-02T15:00', duration is 45 minutes as an integer, both attendees are in an array, and title 'Design review' is sensible and matches the user request.

GPT-5.4 Nano

openai/gpt-5.4-nano

10/10
top judge scorecheapest scoring 8+
Latency
326ms
Cost
$0.00015
Per 1,000
$0.14

152 in · 92 out (43 reasoning) · 2 words · checks 7/7

  • json_valid
  • json_keys
  • regex
  • contains
  • regex
  • contains
  • contains
{"name":"create_event","arguments":{"title":"Design Review","start":"2026-10-02T15:00","duration_minutes":45,"attendees":["[email protected]","[email protected]"]}}

Judge: All parameters correctly formatted: tool name correct, ISO 8601 start datetime accurate, duration is integer 45, both attendees in array, and title is sensible and appropriate.

GPT-5.3-Codex

openai/gpt-5.3-codex

10/10
Latency
405ms
Cost
$0.00169
Per 1,000
$1.69

152 in · 102 out (53 reasoning) · 2 words · checks 7/7

  • json_valid
  • json_keys
  • regex
  • contains
  • regex
  • contains
  • contains
{"name":"create_event","arguments":{"title":"Design review","start":"2026-10-02T15:00","duration_minutes":45,"attendees":["[email protected]","[email protected]"]}}

Judge: All parameters are correct: tool name is 'create_event', ISO start datetime is '2026-10-02T15:00', duration is 45 minutes as an integer, both attendees are in an array, and title 'Design review' is sensible and matches the user request.

Frequently asked

What does this prompt test?

Tool call: Correct tool name, ISO start 2026-10-02T15:00, integer duration 45, both attendees in an array, a sensible title. The deterministic checks are json_valid, json_keys, regex, contains, regex, contains, contains.

Which model should I pick for this task?

If the judge's bar of 8/10 is good enough for you, GPT-5.4 Nano at $0.14 per 1,000 runs. If you need the top score, GPT-5.4 Nano at $0.14 per 1,000 runs.

If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.

Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.

AI cost intelligence

Stop overpaying for AI tools

Join the OverpayingForAI list for pricing updates, cheaper alternatives, and practical buying guidance.

Now tracking 50+ AI tools, models, platforms, subscriptions, coding tools, and automation products.

We use your email only for OverpayingForAI updates. Unsubscribe anytime.