Qwen3.8 Max (0902) vs Qwen3.8 27B vs Qwen3.8 Flash vs Qwen3.7 Plus on people mentioned to a json array
Alibaba models side by side on "People mentioned to a JSON array": Qwen3.8 Flash scores 9/10; Qwen3.8 Flash is the cheapest answer scoring 8+ at $0.11 per 1,000 runs. Outputs, checks, judge reasons, latency and cost.
The prompt every model received
System
You are a data extraction engine. Reply with JSON only, no prose, no code fences.
User
From the text, return a JSON array of every person mentioned. Each element is an object with keys name, role, email (use null when an email is not given). Preserve the order in which people first appear.
Text: The migration is led by Dana Okafor (Head of Platform, [email protected]). She is supported by Luis Ferreira, a staff engineer, whose address is [email protected]. Questions about billing go to Mei Tanaka, our finance lead; she prefers phone calls and has not shared an email.
Rubric for the judge: A three-element JSON array in order (Dana, Luis, Mei) with correct roles, emails and a JSON null for Mei.
Side by side
Every cell is one OpenRouter call at temperature 0 with the prompt's token cap and reasoning effort "low" where the model supports it. Cost is usage × the catalogue rate in models.json. Quality is one judge call to anthropic/claude-haiku-4.5 against the prompt's rubric, cached per prompt version.
Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902
9/10
Latency
1.1s
Cost
$0.00121
Per 1,000
$1.21
172 in · 144 out (81 reasoning) · 8 words · checks 6/6
✓json_valid
✓regex
✓contains
✓contains
✓contains
✓regex
[{"name":"Dana Okafor","role":"Head of Platform","email":"[email protected]"},{"name":"Luis Ferreira","role":"staff engineer","email":"[email protected]"},{"name":"Mei Tanaka","role":"finance lead","email":null}]
Judge: Correct structure, order, and content with all three people and proper null for Mei's email, but 'staff engineer' and 'finance lead' lack capitalization consistency with 'Head of Platform'.
Qwen3.8 27B
qwen/qwen3.8-27b
9/10
Latency
449ms
Cost
$0.00048
Per 1,000
$0.48
172 in · 175 out (112 reasoning) · 8 words · checks 6/6
✓json_valid
✓regex
✓contains
✓contains
✓contains
✓regex
[{"name":"Dana Okafor","role":"Head of Platform","email":"[email protected]"},{"name":"Luis Ferreira","role":"staff engineer","email":"[email protected]"},{"name":"Mei Tanaka","role":"finance lead","email":null}]
Judge: Correct structure, order, and content with all three people and their details accurately captured; minor deduction only for inconsistent capitalization of 'staff engineer' vs 'Head of Platform' and 'finance lead'.
Qwen3.8 Flash
qwen/qwen3.8-flash
9/10
top judge scorecheapest scoring 8+
Latency
220ms
Cost
$0.00011
Per 1,000
$0.11
146 in · 186 out (126 reasoning) · 8 words · checks 6/6
✓json_valid
✓regex
✓contains
✓contains
✓contains
✓regex
[{"name":"Dana Okafor","role":"Head of Platform","email":"[email protected]"},{"name":"Luis Ferreira","role":"staff engineer","email":"[email protected]"},{"name":"Mei Tanaka","role":"finance lead","email":null}]
Judge: Correct structure, order, and content with all three people and their details accurately captured; minor deduction only for inconsistent capitalization of 'staff engineer' vs 'Head of Platform' and 'finance lead'.
Qwen3.7 Plus
qwen/qwen3.7-plus
9/10
Latency
1.8s
Cost
$0.00119
Per 1,000
$1.19
146 in · 892 out (776 reasoning) · 33 words · checks 6/6
Judge: Correct structure, order, and content with all three people and proper null handling, but 'staff engineer' and 'finance lead' lack capitalization consistency with 'Head of Platform'.
Frequently asked
▸What does this prompt test?
Extract JSON: A three-element JSON array in order (Dana, Luis, Mei) with correct roles, emails and a JSON null for Mei. The deterministic checks are json_valid, regex, contains, contains, contains, regex.
▸Which model should I pick for this task?
If the judge's bar of 8/10 is good enough for you, Qwen3.8 Flash at $0.11 per 1,000 runs. If you need the top score, Qwen3.8 Flash at $0.11 per 1,000 runs.
If our calculators helped you cut down on hidden AI wallet leaks, thanks for using them. A tiny fraction of your savings is what keeps our pricing indexes updated daily.
Not sure which AI is cheapest for your use case? Find out in 30 seconds — no signup required.