TokenOptimal
Expected cost per accepted outcome

The cheap model
was not cheap.

You started on the cheap one. It failed. You retried. It failed again. You moved to the big model and it worked. Your bill says the cheap model costs four cents.

Choosing it cost ninety. Every benchmark, every pricing page and every billing export tells you the four. This site is about the ninety.

One decision, three attempts

code.fix-diagnose
01cheap model rejected$0.04
02cheap model, retry rejected$0.05
03frontier model accepted$0.61
your time, two failures at $0.10 $0.20
per-attempt price of the cheap model $0.04
what starting cheap actually cost $0.90

Both numbers are true. Only one of them is a decision. A failed attempt is charged back to the model you chose first — that is the whole method.

The one number

Not cost per attempt. If you start a task on model A and it fails, everything you spend afterwards — a retry, an escalation to model B — is charged back to the decision to start with A. A model's ECAO is what it actually costs to get a result you accept, including its failures. Each rejected attempt is also charged a flat $0.10 for your own time, because without that a model that fails 70% of the time at a fifth of the price comes out "cheapest" — arithmetically true, practically wrong.

List prices, always

Every number here is list-USD from the vendor price sheet. On a subscription plan you never pay list — but quota is metered in roughly the same proportion, so minimising list-USD is the same act as minimising plan burn.

Chains, not attempts

A run that rescued an earlier run points at it, and the whole chain is charged to the first decision. If nobody logs the follow-up, the chain is just its root and ECAO degrades to attempt cost rather than breaking.

Never estimated

Token counts come out of the harness log — Claude Code's session files, Codex's rollouts. An estimated token count poisons a ranking invisibly, because the output still looks like a number.

Plan windows, measured

Nobody publishes how much work a five-hour window holds. It is measurable: if turns costing X list-USD moved a window by P points, the window holds about X / (P/100). Codex writes its own percentage into every turn, so that plan calibrates itself from history. Claude Code logs nothing — but the endpoint behind its /usage command answers with the same two numbers.

Stated bias

These are measurements with a known bias, not published facts. Work done outside the harness — a claude.ai chat, the ChatGPT web app, another machine on the same account — moves the window but never reaches the logs. Every capacity here is therefore a floor, and it is reported with its run count and its date, or it is not reported at all.

The guide

One section per kind of work. A row carrying a tier is evidence — real runs, real token counts, a human verdict. A block marked opinion is a placeholder standing in until evidence displaces it, and it is labelled that way because pretending otherwise is how a guide becomes folklore.

Proven8+ rated, 2+ raters
Measured5+ rated
Anecdotal1 to 4 rated
Extrapolatedfamily evidence only
Opinionnothing rated yet

Tier measures how much evidence exists, not how good a model is. A model can be proven bad.

How a run becomes a ranking

A run is one task, one model, one effort. A human says accepted or rejected, optionally 1–5, optionally "this one rescued that one". All runs for one (task, model, effort) become a cell: first-try success with an 80% Wilson interval — Wilson because n is small and the rate sits near 0 or 1 exactly when it matters — shrunk toward the family's pooled rate, so one anecdote moves the estimate but cannot set it to certainty.

What this is not, yet

  • Single-rater. Everything above measured needs a second rater. The schema carries a rater on every record from day one, so pooled evidence across operators needs no migration — but the pooling does not exist yet.
  • Claude Code does not log reasoning effort. Its runs pool into default unless the person rating supplies one. A Claude effort comparison here is not a measured comparison.
  • Some prices are second-hand. Rows marked secondary came from aggregator sites and have not been checked against the vendor's own page. Do not quote one without checking it.
  • A run is a whole turn. A turn that mixed two kinds of work is labelled as the one it mostly was.
List prices on file