This is the answer to the question everyone asks after two weeks of use: "how do I make it know my documents?" Your prices, your procedures, your meeting notes, your product documentation. None of that appears in what a model learned.
Three solutions exist in theory. Pasting the documents into every conversation, which does not hold beyond a few pages. Retraining the model, which is expensive and unsuited to information that changes. Or automatically retrieving the right passages at question time: that is RAG, and it is the right answer in almost every case.
Definition
RAG, for retrieval augmented generation, is a method that automatically searches a document base for relevant passages, then supplies them to the model along with the question, so that it answers from those passages rather than from memory.
The mechanism has four steps. Your documents are cut into chunks. Each chunk is turned into an Embedding, a numerical representation of its meaning. When a question is asked, the system finds the chunks closest in meaning to the question. Those chunks are slipped into the Prompt, and the model writes its answer from them.
The search is on meaning, not words. A question about "the refund period" finds a paragraph about "return within fifteen days", even though no word matches. That is what separates RAG from a plain internal search engine.
Why it is almost always the right choice
| Approach | Cost | Updating | Traceable |
|---|---|---|---|
| Paste everything into the conversation | High and repeated | Immediate | Yes |
| Fine-tuning | High to set up | New training run | No |
| RAG | Low per question | Immediate | Yes |
The "traceable" column is what decides it for professional use. A RAG system can cite the document each claim came from. A retrained model cannot: the information is diluted across its parameters, with no trace of origin.
The "updating" column matters just as much. Your prices change, your procedures evolve. With RAG, replacing the document is enough. With fine-tuning, training has to start again.
What makes a RAG good
The chunking
This is the setting that decides everything and nobody talks about it. Chunks that are too short lose context, chunks that are too long drown the useful information. Good chunking follows the structure of the document, by section rather than by character count.
The quality of the source
A RAG plugged into outdated documentation answers in an outdated way, confidently. Cleaning up the documents is a prerequisite, not a later improvement.
The number of chunks retrieved
Too few and the answer is incomplete. Too many and the useful information dilutes inside the Context window. Between three and five suits most uses.
The fallback instruction
Without an explicit instruction such as "if the answer is not in the extracts, say so", the model will fill gaps with its general knowledge. You then get an indistinguishable mix of your documents and a Hallucination.
Where to put it in place
The cases that repay the effort look alike: a stable corpus, repetitive questions, and a real cost to a wrong answer.
First-line customer support comes top. Then access to internal procedures, where "how do we do this again" comes up every week. Then product documentation, where volume always exceeds what a person retains.
A badly tuned RAG is more dangerous than no RAG. It produces answers that appear grounded in your documents when they come from elsewhere. Always require the extracts used to be displayed: without that traceability you can verify nothing.
Why a RAG answers badly, and where to start
A system giving poor answers almost always has one of the five causes below. The symptom points straight to the right one, which saves rebuilding everything.
| Symptom | Most likely cause | Fix |
|---|---|---|
| Correct but incomplete answers | Too few extracts retrieved | Move from three to five extracts |
| Off-topic answers | Chunking too fine, context lost | Chunk by section rather than character count |
| It invents despite the base | No fallback instruction | Add "if it is not in the extracts, say so" |
| Outdated information | Documents not kept current | Clean the base, not the settings |
| Good on some questions, bad on others | A subject missing from the documentation | Write the missing page |
The last line is the most frequent and the most misdiagnosed. People spend hours tuning chunking and retrieval when the information exists nowhere in the base. Before any technical adjustment, check that a person could find the answer by reading your documents.
Frequently asked questions
Do you need technical skills to build one?
Less than before. Several automation tools now offer ready-made blocks where dropping in documents is enough to get a working system. Fine-tuning the chunking and selecting the sources remain real work.
Do my documents go to the model provider?
The selected extracts travel in the prompt, yes. The base itself stays where you host it. For sensitive documents, check your provider's retention commitments, or move to a model hosted on your own infrastructure.
Will large context windows make RAG pointless?
No, for economic reasons. Even if your ten thousand pages fitted in the window, sending them with every question would be absurd. RAG does not exist only to work around a technical limit, it exists to send only what is useful.
How do you start without launching a project?
By picking one clearly bounded corpus, for instance your twenty standard support replies, rather than all your documentation. Our n8n course shows that build end to end, from dropping in the documents to checking the extracts cited.