This is the word people reach for when they want "a custom AI". In practice it is almost always the wrong answer to the question being asked, and it is worth knowing why before committing a budget.
Definition
Fine-tuning means continuing the training of an existing model on a set of your own examples, in order to durably change its behaviour.
It is not about building a model from scratch, which would demand resources out of reach. It is about adjusting an already trained model so it adopts a particular way of working: your output format, your style, your trade vocabulary.
A decisive distinction: fine-tuning teaches a way of doing, not information. For a model to know your prices, RAG is the right method. For it to consistently write like your brand, fine-tuning can be justified.
Fine-tuning or RAG: the table that decides
| Your need | The right method |
|---|---|
| Answer from your documents | RAG |
| Respect a very strict output format | Fine-tuning |
| Adopt your brand's tone | System prompt, then fine-tuning if insufficient |
| Know information that changes | RAG |
| Handle very specific trade vocabulary | Fine-tuning |
| Cut the cost of a repetitive task | Fine-tuning a small model |
That last case is the most underrated. A small model specialised on a precise task can match a large generalist model on that task, at a fraction of the cost per call. When volume is high, the upfront investment pays back.
What it actually requires
Examples, in quantity and quality. Count on several hundred representative request-answer pairs. Mediocre examples produce a mediocre model, with remarkable consistency.
A separate test set, in other words an Evaluation. Without examples held back for evaluation, you will not know whether the result improved or whether you simply memorised your examples.
A commitment over time. A specialised model has to be redone whenever your need shifts noticeably, and re-evaluated with every new version of the base model.
The order is always the same: try a better prompt first, then examples inside the prompt, then RAG, and only then fine-tuning. The vast majority of projects stop before the last step, for an equivalent result and far less work.
Frequently asked questions
Does my training data stay private?
Serious providers commit to not reusing it for other customers, and the resulting model is reserved for you. That is verified in the contract rather than on the sales page, and the question deserves to be asked before sending anything.
Does a specialised model lose its other capabilities?
It can, yes. It is a known effect: over-specialised on one task, a model degrades elsewhere. That is no problem if you use it only for that task, and becomes one if you were counting on reusing it everywhere.
Can fine-tuning and RAG be combined?
Yes, and it is often the best configuration for demanding use: fine-tuning fixes the form, RAG supplies up-to-date substance. Each solves a problem the other cannot.
How do you know your need justifies it?
By exhausting the cheaper steps first and measuring each time. If a worked prompt and a well-tuned RAG already get you to the result, the question no longer arises. Our Claude Cowork course covers that progression, from the prompt to the point where going further becomes legitimate.