Cost per token: understanding and controlling your AI bill

Cost per token is the unit price charged by AI providers, different for input and for output.
3 min read
Believemy logo

It is the only unit that lets you compare two providers and forecast a bill. Without it, you choose a model by its name and discover the total at the end of the month.


Definition

Cost per token is the unit price applied to each Token processed by a model. It is usually quoted per million tokens, and it differs by direction.

DirectionWhat is countedRelative price
InputInstruction, history, documents sentThe cheaper side
OutputWhat the model writesThree to five times more
Cached inputInstruction already sent recentlyUp to ten times less

The third line is the one most people ignore, and often the most profitable: see Prompt caching.

Good to know

Never compare two models on headline price alone. A model twice as expensive that answers correctly first time costs less than a cheap one you have to rerun three times. The right comparison unit is cost per successful task.


Estimating a bill before starting

The calculation takes three numbers: calls per month, average input size, average output size.

A concrete example. A support assistant handles 500 requests a month. Each request sends a 500-token instruction, the history and three documentation extracts, around 3,000 tokens of input. The answer runs to 300 tokens. You consume 1.5 million input tokens and 150,000 output tokens a month. All that remains is multiplying by the rates of the model you have in mind.

That five-minute calculation avoids most nasty surprises, and above all shows where to act: here, input weighs ten times more than output, so it is the documentation sent that must shrink, not the length of the answers.

Estimate your monthly bill
500
2,000
200
Model
Estimated monthly cost
6 $
of which 4 $ input · 2 $ output
1,466,667 tokens per month

Indicative rates per million tokens, check with your provider. The calculation converts words to tokens at a ratio of 4 to 3.


The five levers, most to least effective

Reduce what you send. Three well-chosen extracts rather than ten pages. Almost always lever number one.

Enable caching when the same instruction is sent on every call.

Choose a smaller model for simple tasks: classify, extract, sort. The price gap runs from one to fifty.

Limit answer length, since output is the expensive side.

Group your processing so the instruction is not resent per item: see Batch processing.

Warning

The most frequent source of spend is none of those five: it is the conversation growing longer. Every message resends the whole history. By the thirtieth exchange you are paying twenty-nine times over for the accumulated context.


Frequently asked questions

Question

How do you track consumption?

Professional interfaces display consumption per call and expose it in their responses. The indicator to watch is not the monthly total but tokens per operation: a total rising at constant activity signals a leak.


Question

Do prices fall over time?

They have fallen a great deal in recent years at constant quality. That is an argument against over-optimising too early: what is expensive today can become negligible within a year.


Question

Does an agent cost more than a simple call?

Markedly, because an Agent loop makes several calls, each carrying the accumulated history. This is why a pass ceiling is a budget setting as much as a technical one.


Question

How do you keep this under control in a project?

By instrumenting the measurement from the first scenario rather than after the first surprising invoice. Our n8n course builds that tracking into the automations themselves.

Related terms

Discover our aI and automation glossary

The vocabulary of artificial intelligence and automation, explained for people who want to use it in their business, not for people who build the models.

Share this article

Want to help us? Share this article on your networks or even better: on your site, in an article or in your newsletter.