The prompt (verbatim, sent to every lane)
Read this receipt and return only a JSON object in exactly this shape:
{"store": string, "date": "YYYY-MM-DD", "items": [{"name": string, "qty": integer, "unit_price": number}], "total": number}
COPPER KETTLE COFFEE — July 5th, 2026 — two flat whites at $4.75 each, one almond croissant $3.80, oat-milk upgrade $0.70. Card total $14.00.
Live runs are switched off right now — replays (recorded 2026-07-12) are the whole show today.
Live means live: 3 models answer this prompt for real, on my dime — reserving $0.0035 for this run against a shared $0.50 daily reservation ceiling. Replays stay free either way.
Going live sets one cookie (bl_sid, 30 days). It holds a random id, not your identity. The five-runs-a-day gate counts both that signed browser session and a one-way hash of the connection address; raw IP addresses never touch storage, and nothing else rides along.
How scoring works
Every lane gets the identical prompt at the same moment. A small script — not an AI judge — checks each answer against the rule below and stamps it CORRECT or WRONG; the line under each lane shows exactly what the check saw. No partial credit.
This challenge's rule
Parse → validate the exact JSON shape (no extra keys) → require the right store, date 2026-07-05, total 14.00, and exactly three items with the right (qty, unit price) pairs. Item names are displayed, not scored — models will vary "flat whites"/"flat white", and scoring names would be a coin-flip technicality. Scoring strips leading/trailing whitespace and one wrapping pair of code fences or quotes before it judges — content over ceremony. Winner rule: correct → cheapest → fastest.
Prices as listed by OpenAI and Anthropic, July 2026. Costs come from provider usage fields × list prices — nothing is invented.
LIVE BUDGET ················ UTC day live runs ··········· ········open budget reserved ········ $0.00 daily cap ··················· $0.50 ──────────────────────────────────── replays ··········· free, unmetered