tokens

Tokens are not commodities

On OckBench, a benchmark built to measure reasoning efficiency, DeepSeek-V4-Pro at high effort scores 84.0% accuracy. GPT-5.5 at medium effort scores within two points of it. To get there, DeepSeek consumes 43,652 tokens per task. GPT-5.5 consumes 4,692. Two models, near-identical results, 9.3x apart in the unit both are priced in.

Now apply the prices. DeepSeek charges $0.87 per million output tokens; OpenAI charges $30. Per token, DeepSeek looks roughly 34 times cheaper. Multiply through the consumption and the tasks come out at about $0.038 against $0.14, less than four times cheaper. The rate card overstated the advantage by nearly a factor of ten. That factor is the token ratio, and it is the one variable no rate card can express.

The rate card alone cannot price this difference. Per-token pricing treats every token as interchangeable: a million of one model's tokens against a million of another's, dollars per unit, compare and choose. That is how commodities are priced, and it works because a barrel of oil from one supplier burns like a barrel from another. Tokens do not work like this. The intelligence you get per token varies by an order of magnitude between models that reach the same answer, so a rate only becomes a cost once it meets a consumption figure, and the consumption figure is the one you cannot know in advance.

The token is the unit of account. The task is the unit of value. Everything that follows is about the gap between them, and about how the pricing model hides it.

Start with how the meter actually works. You pay for input tokens, the prompt and context you send, and output tokens, the text the model generates, with output typically priced four to six times higher. Reasoning models added a third category that most rate cards do not list separately: the internal chain of thought a model generates before producing its visible answer. Across OpenAI, Google, and Anthropic, these reasoning tokens are billed as output tokens. In OpenAI's accounting they are literally a subset of the completion token count, reported in a usage detail field but part of the same billed pool. Google bills the full thought tokens the model generates even when the API surfaces only a summary. Anthropic states plainly that the billed output count will not match the visible token count in the response.

Here is the correction most people need: whether you see the reasoning trace is a display decision, not a billing decision. DeepSeek returns the raw chain. OpenAI returns a summary. Anthropic returns a condensed version on recent models. In every case you pay for the full internal chain, including, in the summarised cases, tokens that nobody at any price gets to read.

The hidden fraction is not small. A response showing 500 visible output tokens can consume 2,000 to 5,000 billed tokensonce reasoning is counted, and a common budgeting heuristic for reasoning models is to multiply expected output by three to five. One developer's post on OpenAI's own forum records a $5 estimate for a million tokens arriving as a $20 charge because the reasoning was metered as output.

And what happens to the reasoning afterwards depends on the provider, and neither path is kind. On OpenAI's default behaviour, reasoning from earlier turns is not carried into the next; the visible answer persists as context and the deliberation is discarded. Gemini can carry thought state forward, but thought signatures sent back with a request are charged as input tokens: keep the reasoning and you pay for it a second time at the ingest rate. Bought as output, then discarded or re-billed. Either way the deliberation never becomes an asset.

So a growing share of the billed unit cannot be read, only counted after the model has generated it, and the readable share is not comparable across models anyway. This is where the benchmark evidence gets structural.

OckBench proposes per-token intelligence as an evaluation axis: a superior model must not only reach the right answer but reach it with minimal token consumption. Its findings are blunt. Models with similar accuracy can differ by more than 26x in average output length. Smaller models often incur what the authors call an Overthinking Tax, verbose reasoning chains that make a cheap model expensive to serve; same-size 7B models with similar accuracy differ 3.3x in tokens and 5x in latency. And the pattern runs in the direction you would not want if you were defending the commodity framing: the authors observe that smarter, larger models generally produce shorter responses, reaching correct answers with denser reasoning.

Density is the word to hold onto. A separate study across 25 models found that accuracy rankings and token-efficiency rankings diverge on a controlled reasoning benchmark, a Spearman correlation of 0.63, and that verbalization overhead, the share of tokens a model spends beyond what the answer needs, varies roughly 9x across models, only weakly related to model scale. Efficiency is not a byproduct of capability. It is a capability. The rate card prices it at zero.

The commodity framing fails at an even lower level: the vendors do not agree on what a token is. Each provider uses its own tokenizer, so identical text splits into different counts; one cross-vendor analysis puts the gap at low single digits for English prose and 10 to 20% for code. A cross-vendor price comparison divides different numerators by different denominators and presents the result as a like-for-like number. It is not one.

None of this makes per-token billing irrational. Tokens track inference work, they are mechanically countable at scale, and a provider needs a meter it can read in microseconds, not after a task review. The mistake is not the meter. The mistake is reading the meter as the economics: a rate per token describes how a service charges, not what a workload will cost, and not whether the result was worth producing.

The market has noticed. Artificial Analysis, whose model index is the closest thing the industry has to a public price-performance ledger, now reports cost per task, computed from the actual input, cached, reasoning, and answer tokens each model consumed across the benchmark workload. Their stated reason concedes the whole argument: models that produce longer answers or more reasoning tokens have a higher cost per task even at identical per-token prices. Vendors are competing on the same axis. Google's launch pitch for its newest Flash tier leads not with accuracy but with consumption, 17% fewer output tokens on the Artificial Analysis Index than its predecessor, and, by Google's own measurement, up to 65% fewer on an agentic coding benchmark.

When the measurement bodies abandon your unit and the vendors start marketing against it, the unit is finished as a basis for comparison. What survives is the meter.

There is an apparent escape from all of this: the flat-fee subscription. A fixed monthly price, no meter, no token arithmetic. It looks like the answer to a broken unit. It is better read as a confession that the unit is broken: providers sell flat fees precisely because tokens are a unit nobody wants to budget in.

But the tokens do not stop being the cost. The variance the buyer no longer sees, the provider now carries, and no provider carries it for long: it comes back as rationing. Usage caps, weekly limits, priority tiers, quiet routing to cheaper models. The meter disappears from the invoice and reappears as a quota you discover by hitting it. The incentive flips as well: at a flat fee, an overthinking model is free to the buyer at the margin, so the efficiency fight moves from the price sheet into the quota design, where only one side can see it.

And the subscription does something worse to the accounting than per-token pricing ever did. Under API pricing, cost per task is at least recoverable after the fact. Under a flat fee, the marginal cost of a task is zero on the invoice and unknown in reality; the fee amortises over consumption that nobody on the buyer's side measures. A flat rate does not remove cost. It hides it, which is the more dangerous condition. Per-token pricing obscures the gap between account and value. The subscription abolishes the ledger.

Which is the real problem. I have argued before that AI cost is a control problem, not a pricing problem: the thing spending the money is an autonomous process with no judgment about its own budget, and the only architectural chokepoint where enforcement can live is the gateway. Token pricing makes that control problem strictly harder. You cannot govern spend denominated in a unit whose value density varies 9x between models, whose consumption is knowable only once it has happened, and whose definition shifts between vendors. Cost per task is the metric that tells the truth, and it has a property that should bother anyone running agents in production: it is only recoverable after the money is spent.

The industry prices ex ante in a unit whose value is knowable only ex post. Every pricing regime with that shape eventually gets replaced by one that charges for the thing the buyer wanted. The question is not whether task-based pricing arrives. The question is who controls the meter when it does.

Unlock the Future of Business with AI

Dive into our immersive workshops and equip your team with the tools and knowledge to lead in the AI era.

Scroll to top