It is the only unit that lets you compare two providers and forecast a bill. Without it, you choose a model by its name and discover the total at the end of the month.
Definition
Cost per token is the unit price applied to each Token processed by a model. It is usually quoted per million tokens, and it differs by direction.
| Direction | What is counted | Relative price |
|---|---|---|
| Input | Instruction, history, documents sent | The cheaper side |
| Output | What the model writes | Three to five times more |
| Cached input | Instruction already sent recently | Up to ten times less |
The third line is the one most people ignore, and often the most profitable: see Prompt caching.
Never compare two models on headline price alone. A model twice as expensive that answers correctly first time costs less than a cheap one you have to rerun three times. The right comparison unit is cost per successful task.
Estimating a bill before starting
The calculation takes three numbers: calls per month, average input size, average output size.
A concrete example. A support assistant handles 500 requests a month. Each request sends a 500-token instruction, the history and three documentation extracts, around 3,000 tokens of input. The answer runs to 300 tokens. You consume 1.5 million input tokens and 150,000 output tokens a month. All that remains is multiplying by the rates of the model you have in mind.
That five-minute calculation avoids most nasty surprises, and above all shows where to act: here, input weighs ten times more than output, so it is the documentation sent that must shrink, not the length of the answers.
Indicative rates per million tokens, check with your provider. The calculation converts words to tokens at a ratio of 4 to 3.
The five levers, most to least effective
Reduce what you send. Three well-chosen extracts rather than ten pages. Almost always lever number one.
Enable caching when the same instruction is sent on every call.
Choose a smaller model for simple tasks: classify, extract, sort. The price gap runs from one to fifty.
Limit answer length, since output is the expensive side.
Group your processing so the instruction is not resent per item: see Batch processing.
The most frequent source of spend is none of those five: it is the conversation growing longer. Every message resends the whole history. By the thirtieth exchange you are paying twenty-nine times over for the accumulated context.
Frequently asked questions
How do you track consumption?
Professional interfaces display consumption per call and expose it in their responses. The indicator to watch is not the monthly total but tokens per operation: a total rising at constant activity signals a leak.
Do prices fall over time?
They have fallen a great deal in recent years at constant quality. That is an argument against over-optimising too early: what is expensive today can become negligible within a year.
Does an agent cost more than a simple call?
Markedly, because an Agent loop makes several calls, each carrying the accumulated history. This is why a pass ceiling is a budget setting as much as a technical one.
How do you keep this under control in a project?
By instrumenting the measurement from the first scenario rather than after the first surprising invoice. Our n8n course builds that tracking into the automations themselves.