Learn why enterprise AI costs rise at rollout, how hidden model calls inflate cost per answer, and how to measure and optimize AI spend.

What Does One AI Answer Actually Cost? Why Enterprise AI Budgets Break At Rollout

Ask the owner of an enterprise AI programme three questions.
How accurate is it? They have an evaluation set, a score, and a chart showing it going up.
How fast is it? There’s a p95 latency figure on a dashboard somewhere.
What does one answer cost? Pause. Then a spreadsheet: a price per million tokens, multiplied by an average prompt length someone estimated in week two of the pilot.
That last number isn’t measured. It’s derived from a price list. And it’s the number the business case was approved on.

Why the price list lies

A price list quotes the cost of a token. Nobody buys tokens. They buy answers, and an answer is not one call.
Trace a single question through a typical production pipeline. The intent gets classified. The query gets rewritten for retrieval. It gets embedded. Candidates come back from a search and get reranked. A tool might be called, its result fed back, and the loop run again. The answer is generated. A guardrail or verification pass checks it. Somewhere in there, a response comes back as malformed JSON and the call is retried.
That’s six to ten model calls behind one reply. The estimate priced one.
Then the inputs grow in ways nobody modelled:
So cost per answer isn’t constant. It’s a distribution, with a long tail, and the tail is where the money goes.

The arithmetic at rollout

Here is how it plays out. The numbers are illustrative, but the shape is common.
Pilot Rollout Yours
Seats 20 1,000
Questions per seat per day 5 8
Working days per month 22 22
Questions per month 2,200 176,000
Cost per question — estimated $0.02 $0.02
Cost per question — measured $0.11 $0.11
Monthly cost — estimated $44 $3,520
Monthly cost — measured $242 $19,360
Annual gap — $190,080
Monthly cost = seats × questions per seat per day × working days × cost per question.
At pilot scale, nobody notices. Forty-four dollars and two hundred and forty-two dollars are both rounding errors on a project budget. The 5.5× gap is invisible until it’s multiplied by a thousand people — at which point a business case built on $42,000 a year is facing $232,000, and it’s facing it in the quarter it was meant to prove itself.
Note the second row, too. Usage per seat goes up at rollout. People who trust a tool ask it more. The success you were hoping for is itself a cost driver.

Where the $0.11 actually went

The more damaging problem isn’t that the estimate was wrong. It’s that when the real bill arrives, nobody can say why it’s wrong, because cost was never attributed to anything smaller than the invoice.
Break one median answer down by stage and the picture usually surprises people:
Stage Cost per answer Share
Intent classification $0.004 4%
Query rewrite $0.006 5%
Embedding + retrieval $0.001 1%
Reranking $0.010 9%
Answer generation $0.055 50%
Verification pass $0.024 22%
Retries $0.010 9%
Total $0.110
Illustrative breakdown.
Generation — the part everyone budgets for — is half. The verification pass someone added in week six to fix a hallucination problem is nearly a quarter. Retries, which appear on no architecture diagram, are another tenth. None of these is visible from a monthly invoice, and none of them can be fixed by someone who can’t see them.

Make cost a first-class metric

Accuracy and latency got their status because teams decided to measure them on every request. Cost needs the same treatment, and it needs three properties to be useful.
Once cost is measured this way, it becomes something you can regress against. A prompt change that adds 2,000 tokens to every request should fail review the way a 300ms latency regression would. It rarely does today, because nobody can see it.

What it unlocks

Measured, attributed cost turns a vague anxiety about “AI spend” into engineering decisions:
You can’t optimise what you priced from a brochure.

The question to ask before rollout

Before the business case goes to the steering committee, ask the team one thing:
“Show me what yesterday’s two-hundredth question cost, stage by stage.”
If they can, you have a cost model. If they can’t, you have a hypothesis — and it will be tested at a thousand seats, in public, at the worst possible moment.

Recent Posts

Learn why enterprise AI costs rise at rollout, how hidden model calls inflate cost per answer, and how to measure and optimize AI spend.
enterprise-AI-access-control
RoRWA