You have probably seen it happen: a long conversation with an assistant, and after a while it forgets an instruction given at the start, or contradicts what it stated twenty messages earlier. That is not a whim, it is the context window being exceeded.
It is one of the most structural limits of generative AI, and one of the least understood. It explains a good share of the disappointments, and mastering it completely changes the quality of what you get.
Definition
The context window is the maximum amount of text a model can take into account at one time. It is measured in Token and covers everything: your instruction, the conversation history, the documents supplied and the answer being written.
Working memory is the most accurate image. It is not a hard drive where the model files information away, it is a desk of fixed size. Everything that must be taken into account has to fit on it at the same time.
An essential nuance: what leaves the window is not forgotten, it never existed for the model. It will not say "I no longer remember", it will answer as if the information had never been given, with the same confidence.
A big window does not mean a good result
Providers advertise ever larger windows, able to swallow hundreds of pages. That is real progress, but it comes with two effects marketing rarely mentions.
Cost is proportional. Filling a large window on every call is paid on every call. A conversation dragging a hundred pages of context costs a hundred pages per message.
Attention gets diluted. An instruction buried in two hundred pages is followed markedly less well than the same instruction with three relevant extracts. Models retain the beginning and the end of what they are given better than the middle.
| Situation | Effective habit |
|---|---|
| One specific document to analyse | Supply it in full |
| An entire document base | Set up RAG |
| A conversation dragging on | Summarise and start fresh |
| A long, recurring instruction | Put it in the System prompt |
One message costs around 300 tokens. The room reserved for the answer is taken off the available total.
Using the available space well
Put what matters at the edges. The main instruction at the start or the end of a message is followed better than in the middle of a long block.
Select rather than pile up. Three well-chosen extracts beat the complete file. Sorting upstream is the best investment you can make in quality.
Structure visibly. Headings, clear separators between the document and the question. A model follows organised text better, exactly like a reader in a hurry.
Watch for the tipping point. When answers start contradicting the beginning of the conversation, the window is saturated. Summarise and restart rather than pushing on.
Frequently asked questions
How do I know I am nearing the limit?
The signs are fairly clear: the model forgets a constraint set at the start, repeats itself, or reopens a decision already made. Professional interfaces usually display context consumption, so you are not flying blind.
Does a larger window make RAG pointless?
No, for reasons of cost more than capacity. Even if your ten thousand pages fitted in the window, sending them with every question would be economically absurd. RAG is not only a workaround for a technical limit, it is a way to send only what is useful.
Does the model keep anything between conversations?
By default, no. Some tools add a memory layer that automatically reinjects elements of previous exchanges, but that is a mechanism built on top of the model, not a capability of the model itself.
How do you organise long work without saturating the window?
By splitting it into stages that each produce a written result, reusable as the starting point of the next. That is a working method more than a technical trick, and our Claude Cowork course applies it to long files, from specification to final report.