You started on the cheap one. It failed. You retried. It failed again. You moved to the big model and it worked. Your bill says the cheap model costs four cents.
Choosing it cost ninety. Every benchmark, every pricing page and every billing export tells you the four. This site is about the ninety.
Both numbers are true. Only one of them is a decision. A failed attempt is charged back to the model you chose first — that is the whole method.
Not cost per attempt. If you start a task on model A and it fails, everything you spend afterwards — a retry, an escalation to model B — is charged back to the decision to start with A. A model's ECAO is what it actually costs to get a result you accept, including its failures. Each rejected attempt is also charged a flat $0.10 for your own time, because without that a model that fails 70% of the time at a fifth of the price comes out "cheapest" — arithmetically true, practically wrong.
Every number here is list-USD from the vendor price sheet. On a subscription plan you never pay list — but quota is metered in roughly the same proportion, so minimising list-USD is the same act as minimising plan burn.
A run that rescued an earlier run points at it, and the whole chain is charged to the first decision. If nobody logs the follow-up, the chain is just its root and ECAO degrades to attempt cost rather than breaking.
Token counts come out of the harness log — Claude Code's session files, Codex's rollouts. An estimated token count poisons a ranking invisibly, because the output still looks like a number.
Nobody publishes how much work a five-hour window holds. It is measurable: if turns costing X list-USD moved a window by P points, the window holds about X / (P/100). Codex writes its own percentage into every turn, so that plan calibrates itself from history. Claude Code logs nothing — but the endpoint behind its /usage command answers with the same two numbers.
These are measurements with a known bias, not published facts. Work done outside the harness — a claude.ai chat, the ChatGPT web app, another machine on the same account — moves the window but never reaches the logs. Every capacity here is therefore a floor, and it is reported with its run count and its date, or it is not reported at all.
One section per kind of work. A row carrying a tier is evidence — real runs, real token counts, a human verdict. A block marked opinion is a placeholder standing in until evidence displaces it, and it is labelled that way because pretending otherwise is how a guide becomes folklore.
Tier measures how much evidence exists, not how good a model is. A model can be proven bad.
A run is one task, one model, one effort. A human says accepted or rejected, optionally 1–5, optionally "this one rescued that one". All runs for one (task, model, effort) become a cell: first-try success with an 80% Wilson interval — Wilson because n is small and the rate sits near 0 or 1 exactly when it matters — shrunk toward the family's pooled rate, so one anecdote moves the estimate but cannot set it to certainty.