Tokens in AI: definition and impact on your bill

A token is the unit an AI model uses to cut up text. It is also the billing unit of every provider.
5 min read
Believemy logo

If you were to keep a single technical term from this whole glossary in order to control your costs, this would be it. The token is the unit in which a model reads, writes and bills. Everything else follows from it.

It is also the word that explains surprise invoices. An automation that seemed to cost a few cents ends up at several hundred euros a month, and the answer is almost always in the number of tokens consumed without anyone noticing.


Definition

A token is a piece of text as the model cuts it up. It is neither a character nor exactly a word: it is an intermediate unit, often a whole short word or a fragment of a long one.

In English, count roughly four tokens for every three words. A 300-word email therefore weighs around 400 tokens, an A4 page about 700, a twenty-page contract around 14,000.

Good to know

Languages are not equal at the cutting stage. English uses fewer tokens than French for the same content, because models were trained mostly on English. For identical content, an instruction written in English often costs 25 to 30 percent less.

How a sentence is cut into tokensLong words are split into several tokens, short words into one.A sentenceAutomation changes everythingAutomation changes everything4 tokens for 3 wordsLong words get split. On average: about 4 tokens for 3 words.


Why this unit decides your bill

Every provider bills by the token, with a distinction many discover too late: input and output are not priced the same.

DirectionWhat it isRelative price
InputEverything you send: instruction, history, documentsThe cheaper side
OutputWhat the model producesOften three to five times more

That asymmetry has a counter-intuitive consequence. Asking for a ten-line summary of a fifty-page document costs little, while asking for a long text from a one-line instruction costs a lot. It is not the question that weighs, it is the answer.


The three traps that blow up the bill

History resent with every message

In a conversation, the entire previous exchange is resent to the model with each new message. By the twentieth message you are paying nineteen times over for the accumulated context. It is the biggest source of waste, and the least visible.

The whole document pasted in

Sending a hundred-page manual to ask about one paragraph works, but costs a hundred pages per question. That is precisely the problem RAG solves.

The loop with no ceiling

An AI agent chaining steps consumes on every pass. With no limit on the number of steps, an agent stuck on an unsolvable problem can burn a monthly budget overnight.

Warning

The right monitoring habit is not watching the total at month end, but the number of tokens per operation. A total that doubles because activity doubled is healthy. A total that doubles at constant activity signals a leak.


How to bring consumption down

Start a fresh conversation. As soon as an exchange drifts from its original subject, open another one with a short summary. That cuts accumulation at the root.

Use caching where it exists. If the same long instruction is sent on every call, most providers let you cache it at a reduced rate. At volume, the saving is substantial.

Match model size to the task. A small model often costs ten times less than a large one. On sorting or extraction, quality is equivalent.

Ask for short answers. Since output is the expensive side, adding "in three sentences" acts directly on the bill.


How many tokens are in your usual content

These orders of magnitude are enough to size a project without a counting tool. They apply to English; French is less economical, closer to five tokens for every three words.

ContentWordsTokens, roughly
A short message3040
An everyday email150200
An A4 page500670
A detailed product description8001,070
A blog article1,5002,000
A twenty-page contract10,00013,300
A two-hundred-page book80,000107,000

Two reference points follow from that table. A thirty-message conversation carries, by the last message, the equivalent of a blog article as input, and you pay for it again with every new message. And a document you paste to ask a single question costs its full weight on every question asked, which becomes absurd beyond three or four.

Cut a sentence into tokens
Automationchangeseverythingforasmallbusiness.
7
words
10
tokens
1.4
tokens per word
51
characters

An estimate, not an exact count. Every model has its own splitting: this simulator reproduces the principle and gives the right order of magnitude, not your provider's figure.


Frequently asked questions

Question

How do I know how many tokens my text contains?

Most providers offer a counting tool, and their professional interfaces display the real consumption of each call. For a quick estimate, the three-tokens-per-four-words rule is more than enough to size a project.


Question

Does a token cost the same everywhere?

No, the gap between models runs from one to fifty. This is why thinking in tokens rather than in euros makes comparison possible: measure how many tokens your use consumes, then apply the rate of the model you are considering.


Question

Does the length of my question affect the quality of the answer?

Up to a point, yes: giving context clearly improves the result. Beyond that, the effect reverses. An instruction buried in ten pages of incidental material is followed less well than a clean instruction with two well-chosen extracts.


Question

How do you keep these costs under control in an automation?

By setting a step ceiling and a budget alert from the start, before the automation ever runs in production. Our n8n course covers that alongside building the scenarios themselves, rather than as an afterthought.

Related terms

Discover our aI and automation glossary

The vocabulary of artificial intelligence and automation, explained for people who want to use it in their business, not for people who build the models.

Share this article

Want to help us? Share this article on your networks or even better: on your site, in an article or in your newsletter.