RAG: making an AI answer from your own documents

RAG lets an AI model answer from your documents by retrieving the useful passages before writing.
5 min read
Believemy logo

This is the answer to the question everyone asks after two weeks of use: "how do I make it know my documents?" Your prices, your procedures, your meeting notes, your product documentation. None of that appears in what a model learned.

Three solutions exist in theory. Pasting the documents into every conversation, which does not hold beyond a few pages. Retraining the model, which is expensive and unsuited to information that changes. Or automatically retrieving the right passages at question time: that is RAG, and it is the right answer in almost every case.


Definition

RAG, for retrieval augmented generation, is a method that automatically searches a document base for relevant passages, then supplies them to the model along with the question, so that it answers from those passages rather than from memory.

The mechanism has four steps. Your documents are cut into chunks. Each chunk is turned into an Embedding, a numerical representation of its meaning. When a question is asked, the system finds the chunks closest in meaning to the question. Those chunks are slipped into the Prompt, and the model writes its answer from them.

Good to know

The search is on meaning, not words. A question about "the refund period" finds a paragraph about "return within fifteen days", even though no word matches. That is what separates RAG from a plain internal search engine.

How RAG worksThe question is used to retrieve extracts from your documents, and the model writes its answer from those extracts.QuestionMeaning-based searchYour documentschunked and indexedRelevant extractsGrounded answerThe model answersonly from the extracts


Why it is almost always the right choice

ApproachCostUpdatingTraceable
Paste everything into the conversationHigh and repeatedImmediateYes
Fine-tuningHigh to set upNew training runNo
RAGLow per questionImmediateYes

The "traceable" column is what decides it for professional use. A RAG system can cite the document each claim came from. A retrained model cannot: the information is diluted across its parameters, with no trace of origin.

The "updating" column matters just as much. Your prices change, your procedures evolve. With RAG, replacing the document is enough. With fine-tuning, training has to start again.


What makes a RAG good

The chunking

This is the setting that decides everything and nobody talks about it. Chunks that are too short lose context, chunks that are too long drown the useful information. Good chunking follows the structure of the document, by section rather than by character count.

The quality of the source

A RAG plugged into outdated documentation answers in an outdated way, confidently. Cleaning up the documents is a prerequisite, not a later improvement.

The number of chunks retrieved

Too few and the answer is incomplete. Too many and the useful information dilutes inside the Context window. Between three and five suits most uses.

The fallback instruction

Without an explicit instruction such as "if the answer is not in the extracts, say so", the model will fill gaps with its general knowledge. You then get an indistinguishable mix of your documents and a Hallucination.


Where to put it in place

The cases that repay the effort look alike: a stable corpus, repetitive questions, and a real cost to a wrong answer.

First-line customer support comes top. Then access to internal procedures, where "how do we do this again" comes up every week. Then product documentation, where volume always exceeds what a person retains.

Warning

A badly tuned RAG is more dangerous than no RAG. It produces answers that appear grounded in your documents when they come from elsewhere. Always require the extracts used to be displayed: without that traceability you can verify nothing.


Why a RAG answers badly, and where to start

A system giving poor answers almost always has one of the five causes below. The symptom points straight to the right one, which saves rebuilding everything.

SymptomMost likely causeFix
Correct but incomplete answersToo few extracts retrievedMove from three to five extracts
Off-topic answersChunking too fine, context lostChunk by section rather than character count
It invents despite the baseNo fallback instructionAdd "if it is not in the extracts, say so"
Outdated informationDocuments not kept currentClean the base, not the settings
Good on some questions, bad on othersA subject missing from the documentationWrite the missing page

The last line is the most frequent and the most misdiagnosed. People spend hours tuning chunking and retrieval when the information exists nowhere in the base. Before any technical adjustment, check that a person could find the answer by reading your documents.


Frequently asked questions

Question

Do you need technical skills to build one?

Less than before. Several automation tools now offer ready-made blocks where dropping in documents is enough to get a working system. Fine-tuning the chunking and selecting the sources remain real work.


Question

Do my documents go to the model provider?

The selected extracts travel in the prompt, yes. The base itself stays where you host it. For sensitive documents, check your provider's retention commitments, or move to a model hosted on your own infrastructure.


Question

Will large context windows make RAG pointless?

No, for economic reasons. Even if your ten thousand pages fitted in the window, sending them with every question would be absurd. RAG does not exist only to work around a technical limit, it exists to send only what is useful.


Question

How do you start without launching a project?

By picking one clearly bounded corpus, for instance your twenty standard support replies, rather than all your documentation. Our n8n course shows that build end to end, from dropping in the documents to checking the extracts cited.

Related terms

Discover our aI and automation glossary

The vocabulary of artificial intelligence and automation, explained for people who want to use it in their business, not for people who build the models.

Share this article

Want to help us? Share this article on your networks or even better: on your site, in an article or in your newsletter.