Ask the owner of an enterprise AI programme three questions.
How accurate is it? They have an evaluation set, a score, and a chart showing it going up.
How fast is it? There’s a p95 latency figure on a dashboard somewhere.
What does one answer cost? Pause. Then a spreadsheet: a price per million tokens, multiplied by an average prompt length someone estimated in week two of the pilot.
That last number isn’t measured. It’s derived from a price list. And it’s the number the business case was approved on.
Why the price list lies
A price list quotes the cost of a token. Nobody buys tokens. They buy answers, and an answer is not one call.
Trace a single question through a typical production pipeline. The intent gets classified. The query gets rewritten for retrieval. It gets embedded. Candidates come back from a search and get reranked. A tool might be called, its result fed back, and the loop run again. The answer is generated. A guardrail or verification pass checks it. Somewhere in there, a response comes back as malformed JSON and the call is retried.
That’s six to ten model calls behind one reply. The estimate priced one.
Then the inputs grow in ways nobody modelled:
- Retrieved context inflates input tokens. The pilot estimate assumed a 1,500-token prompt. In production, the system prompt is 2,000 tokens, the retrieved passages are 6,000, and the input is several times what was budgeted before a word of output is written.
- Conversation history compounds. The sixth turn of a session carries the previous five with it. Cost per question rises with every follow-up, which is exactly how engaged users behave.
- Agentic loops are variable. A question that resolves in two iterations today takes seven tomorrow because a document was updated. Same question, three times the bill.
- Output is priced higher than input, and verbose narration is the default unless someone constrains it.
So cost per answer isn’t constant. It’s a distribution, with a long tail, and the tail is where the money goes.
The arithmetic at rollout
Here is how it plays out. The numbers are illustrative, but the shape is common.
| Pilot | Rollout | Yours | |
|---|---|---|---|
| Seats | 20 | 1,000 | |
| Questions per seat per day | 5 | 8 | |
| Working days per month | 22 | 22 | |
| Questions per month | 2,200 | 176,000 | |
| Cost per question — estimated | $0.02 | $0.02 | |
| Cost per question — measured | $0.11 | $0.11 | |
| Monthly cost — estimated | $44 | $3,520 | |
| Monthly cost — measured | $242 | $19,360 | |
| Annual gap | — | $190,080 |
Monthly cost = seats × questions per seat per day × working days × cost per question.
At pilot scale, nobody notices. Forty-four dollars and two hundred and forty-two dollars are both rounding errors on a project budget. The 5.5× gap is invisible until it’s multiplied by a thousand people — at which point a business case built on $42,000 a year is facing $232,000, and it’s facing it in the quarter it was meant to prove itself.
Note the second row, too. Usage per seat goes up at rollout. People who trust a tool ask it more. The success you were hoping for is itself a cost driver.
Where the $0.11 actually went
The more damaging problem isn’t that the estimate was wrong. It’s that when the real bill arrives, nobody can say why it’s wrong, because cost was never attributed to anything smaller than the invoice.
Break one median answer down by stage and the picture usually surprises people:
| Stage | Cost per answer | Share |
|---|---|---|
| Intent classification | $0.004 | 4% |
| Query rewrite | $0.006 | 5% |
| Embedding + retrieval | $0.001 | 1% |
| Reranking | $0.010 | 9% |
| Answer generation | $0.055 | 50% |
| Verification pass | $0.024 | 22% |
| Retries | $0.010 | 9% |
| Total | $0.110 |
Illustrative breakdown.
Generation — the part everyone budgets for — is half. The verification pass someone added in week six to fix a hallucination problem is nearly a quarter. Retries, which appear on no architecture diagram, are another tenth. None of these is visible from a monthly invoice, and none of them can be fixed by someone who can’t see them.
Make cost a first-class metric
Accuracy and latency got their status because teams decided to measure them on every request. Cost needs the same treatment, and it needs three properties to be useful.
- Read back, not derived. Every model call returns its own usage — input tokens, output tokens, cached tokens. Record it from the response, price it at the rate in effect on that day, and store it against the request. A cost that is computed from a spreadsheet is a forecast. A cost read back from the call is a fact.
- Attributed to the stage that spent it. Tag every call with the pipeline stage that made it. Without this, the only lever you have is "use a cheaper model everywhere", which trades accuracy for budget blindly. With it, you can see that the reranker is overspending on easy queries, or that history is ballooning after turn four, and fix that one thing.
- Visible per answer and per session. Put cost next to latency in the trace of every response, and report median and p95, never just the mean. The mean hides the tail. The tail is the seven-iteration agent loop that cost forty times the median and happens a hundred times a day.
Once cost is measured this way, it becomes something you can regress against. A prompt change that adds 2,000 tokens to every request should fail review the way a 300ms latency regression would. It rarely does today, because nobody can see it.
What it unlocks
Measured, attributed cost turns a vague anxiety about “AI spend” into engineering decisions:
- Route classification to a small model, and keep the large one for judgement.
- Replace a model call with ordinary code wherever the job is deterministic. Arithmetic, rules and lookups don't need a language model, and every call removed is cost, latency and a failure mode removed at once.
- Cache what repeats. Summarise history instead of resending it. Cap agent iterations with a measured budget rather than a guess.
You can’t optimise what you priced from a brochure.
The question to ask before rollout
Before the business case goes to the steering committee, ask the team one thing:
“Show me what yesterday’s two-hundredth question cost, stage by stage.”
If they can, you have a cost model. If they can’t, you have a hypothesis — and it will be tested at a thousand seats, in public, at the worst possible moment.