Behind every conversational assistant you use sits a large language model. It is the basic building block of almost everything happening in Generative artificial intelligence, and the term comes up in every conversation without anyone often saying what it means.
Understanding how it works is not a theoretical exercise. It explains why it is excellent at some tasks, disastrous at others, and why the bill climbs as a conversation gets longer.
Definition
A large language model, usually shortened to LLM, is a system trained on enormous volumes of text to predict what comes next.
Its fundamental task is disarmingly simple: you give it a beginning, it proposes the most likely continuation. Repeat that a few hundred times and you have a paragraph. That is all. There is no knowledge base being consulted, no reasoning in the human sense, no verification.
What makes it remarkable is scale. Fed enough text, the model ends up capturing grammar, style, factual knowledge, patterns of reasoning and even the conventions of a sales email. Capabilities nobody programmed explicitly emerge from it.
One practical consequence: the model does not know when it does not know. It will always produce a plausible continuation, including when the right answer would have been "I do not have that information". This is the direct source of Hallucination.
What matters about how it works
It works in tokens, not words
Text is cut into Token before processing. That unit is what measures consumption and therefore what gets billed.
Its working memory is limited
Everything it can take into account at a given moment fits inside its Context window. Beyond that, the oldest information falls out of range. A model does not forget the way a person forgets: what leaves the window never existed for it.
It retains nothing between conversations
Unless a mechanism is added on top, every new conversation starts from nothing. The feeling of continuity comes from the history being resent in full with each message, not from memory in the model.
Its knowledge stops at a date
Training has an end. Anything after that date is unknown to it, unless you supply it in the conversation or it has a search tool.
Choosing a model without getting it wrong
Every provider offers several sizes, and reaching for the most powerful is often the wrong call.
| Need | Right model | Why |
|---|---|---|
| Classify, sort, extract | Small | The task is simple and repeated: unit price dominates |
| Draft, rewrite | Mid-range | Good balance of quality and cost |
| Analyse, decide, code | Large or Reasoning model | A mistake costs more than the price difference |
The logic is the same as hiring: you do not hand invoice entry to a finance director, and you do not hand strategy to an intern. Start small and move up only when quality falls short.
Best practices
Supply context rather than relying on general knowledge. A model given your document answers far better than a model asked to guess what your company does.
Do not ask it to count or calculate. It produces plausible continuations, not exact results. For arithmetic, give it a tool or do it yourself.
Check anything that looks like a precise fact. A name, a date, a figure, a legal reference. The more specific the detail, the more it deserves verification.
Cut conversations that drag on. A hundred-message thread is expensive and dilutes the instruction. Start fresh with a summary rather than accumulating.
Frequently asked questions
Are LLM and generative AI the same thing?
No. Generative AI is the category, the LLM is the piece specialised in text. An image generator is generative AI without being a language model.
Why are two answers to the same question different?
Because the model picks among several likely continuations, with an amount of randomness set by the Temperature. That variability is an asset when you want ideas, a flaw when you want reproducible processing. Most professional interfaces let you reduce it.
Can an LLM work on my internal documents?
Yes, in two ways. Either you paste them into the conversation, which is enough for a one-off document. Or you set up a RAG system, which automatically retrieves the useful passages from your document base. The second becomes necessary as soon as the volume exceeds what a conversation can hold.
How do you learn to get real time savings from it?
By working on your own files rather than generic examples: a model becomes useful when it is given real context. Our Claude Cowork course is built on that principle, with cases drawn from the daily life of a small business.