This is the most profitable and least known optimisation. It requires no change to your instructions or your results, only a setting. On regular use it commonly halves or thirds the bill.
Definition
Prompt caching is a mechanism letting the provider temporarily keep a portion of already-processed context, and bill it at a heavily reduced rate on subsequent calls.
The principle is simple: if every call sends the same 3,000 tokens of instruction and documentation, the model does not need to reprocess them in full each time.
The condition is structural: caching works on the beginning of the prompt. So put what does not change first, and what changes at the end. A single variable at the start of the instruction is enough to invalidate the whole cache.
What to cache
| Put first, cached | Leave at the end |
|---|---|
| The System prompt | The user's question |
| Reference documentation | Today's data |
| Style examples | Recent history |
| Business rules | Variable context |
A frequent mistake is inserting today's date or a session identifier at the head of the instruction. It looks harmless and it cancels the whole benefit: every call becomes a first call.
When it is worth it
An assistant with a long fixed instruction. The ideal case: the instruction is identical on every call, only the question changes.
A RAG system with recurring extracts. If the same documents come back often, they are cache candidates.
An AI agent that loops. Every pass resends the whole history. Caching comes into its own there, since the beginning never changes.
Conversely, if each of your calls is different and short, caching brings nothing and can even cost slightly more to write.
Cache lifetime is short, from a few minutes to a few tens of minutes depending on the provider. On very spread-out use it expires between calls and serves no purpose. A regular rhythm is needed for it to pay.
Frequently asked questions
Do you have to enable it manually?
It depends on the provider: some apply it automatically, others require explicitly marking the portion to keep. Worth checking in the documentation, because the gain is too large to leave to chance.
Does the result change?
No, it is a billing and performance mechanism. The answer produced is the same, it simply arrives faster and costs less.
Is my data stored?
The cache is temporary and isolated per account. It does not constitute retention of your data in the contractual sense, but if the subject is sensitive the question is worth putting to your provider.
How do you know your cache is working?
API responses report the number of tokens read from cache. If that number stays at zero, your instruction varies at the head. Our n8n course shows how to structure calls to benefit from it.