Clef · Try the live demo →
Glossary

RAG (Retrieval-Augmented Generation)

RAG is an architecture in which a language model looks something up in a defined body of documents before answering, and cites where it found the answer. The difference from relying on model knowledge is traceability: each answer can be checked against the cited passage in the original. For most enterprise work it is the default design, because it makes current internal information usable without touching the model.

How does it actually work?

Three steps. The question is turned into a numeric representation and matched against a prepared index of your documents. The passages that match are handed to the model as context. The model answers from those passages and cites where each statement came from.

The quality lever is almost never the model. It is how documents were split, what metadata travelled with each chunk, and how retrieval was ranked. Two systems on the same model can differ enormously in usefulness because one of them knows which version of a document is current and the other does not.

Why does it matter legally?

Because it changes what happens to your data. In a RAG design the documents stay where they are; only the passages relevant to a question travel to the model, and only for the duration of that answer. Nothing is written into model weights.

That distinction is what makes the GDPR analysis tractable. Training on personal data raises purpose-limitation questions under Article 5(1)(b) that most projects cannot answer. Looking something up and discarding it afterwards raises a much narrower set — which is why we start here rather than with fine-tuning.

Where does RAG fail?

When the underlying documents contradict each other and nothing records which version governs. The system will answer correctly from a superseded policy, cite it properly, and be wrong. This is a content problem wearing a technology costume, and no model upgrade fixes it.

The second failure is the question that nobody wrote down anywhere. If the answer does not exist in the corpus, a well-built system says so. If it does not say so, it was built to always answer — and that is the setting to check before anything else.

What it is not to be confused with

Fine-tuning

Fine-tuning changes the model; RAG changes what the model is shown. For "the system should know our documents", RAG is almost always the answer — it is cheaper, updates the moment a document changes, and produces citations. Fine-tuning earns its cost for style and output format, not for knowledge.

A long context window

Pasting an entire handbook into the prompt works until the handbook is large, expensive to send on every question, or changes weekly. Retrieval solves cost and freshness; a large window does not.

Frequently asked

What does RAG stand for?+

Retrieval-Augmented Generation. Retrieval is the lookup in your own corpus, generation is the model writing the answer, and augmented means the second is constrained by the first. The name describes the order of operations, which is the useful part of it.

Does RAG stop hallucinations?+

It reduces them rather than ruling them out. When a model answers from passages it was given and must cite them, the main cause of fabrication is gone — but a model can still misread or misquote a passage, so the cited place is worth checking. What also has to be configured is permission to say "not found" — a system that must always produce an answer will produce one.

How large does a document corpus need to be?+

Small corpora work fine; the threshold is not size but ambiguity. A hundred well-maintained documents outperform ten thousand where three versions of the same policy sit side by side with no indication which one applies.

Can this run without sending data outside the EEA?+

Yes, and it is the common reason organisations choose this design. Documents stay in your systems, retrieval runs where you place it, and inference can run on hardware inside the EEA. Whether it does is a procurement decision, not a property of RAG.

The question that settles it in ten minutes

Send one typical document and three questions your people actually ask about it. What comes back — with or without citations, with or without an honest "not in here" — tells you more than any architecture diagram.

Request a test