Temperature: tuning how random a model gets

Temperature controls how much a model's answers vary: low for consistency, high for diversity.
3 min read
Believemy logo

This is the setting that explains why the same question asked twice does not give the same answer. Many people discover it while trying to work out why their automation works every other day.


Definition

Temperature is a setting controlling how much randomness goes into choosing the next word. At each step the model has several possible continuations with their probabilities: temperature decides whether it always takes the most likely one or allows itself the others.

ValueBehaviourUse
0Always takes the most likely continuationClassify, extract, decide
0.2 to 0.5Very consistent, slight variationAnswering from documents
0.7 to 1Varied, sometimes surprisingDrafting, looking for ideas
Above 1Goes in all directionsAlmost never useful
Good to know

The rule that avoids the most trouble: low for anything you automate, high for anything you review. An automation must produce the same result on the same input; a draft benefits from offering you something other than the most expected phrasing.


What temperature does not do

It does not make the model more reliable. At zero it always takes the most likely continuation, which does not mean the truest. A Hallucination produced at zero will be reproduced identically on every call.

It does not replace an instruction. If answers go in all directions, the cause is usually an imprecise Prompt, not a setting that is too high.

It does not guarantee reproducibility. Even at zero, two identical calls can differ slightly depending on the provider's infrastructure. It is rare, but enough not to build on an assumption of perfect equality.


The case that comes up most

An automation asks the model to sort incoming requests into three categories. It works well in testing, then starts producing unexpected categories, sometimes with different capitalisation, sometimes an invented fourth category.

Two fixes, in this order. Drop the temperature to zero, which removes the variability. Then ask for Structured output, which constrains the shape instead of suggesting it. The setting alone reduces the problem, structured output removes it.

Warning

Many consumer interfaces do not expose this setting and apply a middle value suited to conversation. If you see troublesome variation in an automation, first check that you actually control it: through an API, yes; in a chat interface, rarely.


Frequently asked questions

Question

What value should you default to?

Zero for any task feeding a Workflow, around 0.7 for drafting. In between there is no magic number: test on your own cases rather than copying a value found online.


Question

Does a low temperature cost less?

No, billing depends on the number of Token, not the setting. That said, a more consistent answer needs fewer retries, which indirectly reduces consumption.


Question

Is it the only setting of its kind?

No, providers offer others acting on the same selection mechanism. In practice, tuning temperature is enough in almost every case, and stacking settings makes behaviour harder to explain.


Question

How do you know your setting is right?

By replaying the same cases several times and seeing whether results vary: that is exactly the role of an Evaluation. Our n8n course covers this setting when connecting a model into a scenario, where consistency becomes a requirement.

Related terms

Discover our aI and automation glossary

The vocabulary of artificial intelligence and automation, explained for people who want to use it in their business, not for people who build the models.

Share this article

Want to help us? Share this article on your networks or even better: on your site, in an article or in your newsletter.