Skip to content
AI Guides

· Updated · 4 min read

What Is RAG? AI Answers From Your Documents

If you've used an AI tool that can answer questions about a PDF you uploaded, or a company's internal knowledge base, you've used RAG — even if you never saw the term. Here's what's actually happening.

The problem RAG solves

A language model's knowledge comes entirely from its training data, which has a cutoff date and obviously can't include your private documents, your company's internal wiki, or anything published after training. Without help, the model simply can't answer questions about that content — it has no way to know it exists.

The two-step trick

Retrieval-Augmented Generation works in two steps, and the name describes exactly what happens:

1. Retrieval. When you ask a question, the system first searches a specific set of documents — using embeddings to find the passages most related to your question by meaning, not just keyword matching — and pulls out the most relevant chunks.

2. Generation. Those retrieved chunks get handed to the language model as extra context, along with your original question, and the model generates its answer based on that specific material — not just its general training.

The result: the model can accurately discuss a document it was never trained on, because the relevant parts are being fed to it fresh, every time, as part of the conversation.

Why this matters more than it sounds

RAG is one of the main ways companies make AI genuinely useful on their own private or recent information, without the expense of retraining a whole model:

  • Customer support tools that answer accurately from a company's actual documentation, not general knowledge.
  • Research assistants that cite specific passages from a set of papers or documents you provide.
  • Internal tools that let employees ask questions in plain English and get answers sourced from company files.

RAG vs fine-tuning: a common mix-up

These two get confused often. Fine-tuning changes the model's underlying behavior through additional training. RAG doesn't touch the model at all — it changes what information the model has access to at the moment it answers. RAG is usually faster to set up, cheaper, and easier to keep up to date (you just update the document set, not retrain anything). See What Is Fine-Tuning? for the full comparison.

The takeaway

Next time an AI tool accurately answers a question about a document you gave it, RAG is very likely the mechanism — search first, then generate an answer grounded in what was found.

See Retrieval-Augmented Generation and Embeddings in our AI Glossary.

Why RAG answers can still be wrong

RAG reduces hallucination considerably. It does not eliminate it, and the ways it fails are worth knowing before you trust a tool built on it.

Retrieval can miss. If the search step does not surface the relevant passage, the model answers from general knowledge instead — often without signalling that it found nothing. An answer built on nothing looks identical to one built on your documents.

Chunking splits meaning. Documents are cut into pieces before being indexed. When a cut lands mid-explanation, the retrieved fragment can be misleading on its own even though the full document was clear.

Contradictory sources produce confident answers. If your document set contains an old policy and a new one, retrieval may return either. The model does not usually know which is current.

Citations prove retrieval, not correctness. A cited passage confirms the system found that text. It does not confirm the summary of it is accurate — and this is the failure people are least likely to catch, because a citation feels like proof.

How to use a RAG tool well

  • Ask questions your documents actually answer. RAG cannot tell you what your sources are missing; it can only work with what is there.
  • Open at least one citation. Ten seconds of checking catches most misreadings and is the single habit that separates careful use from credulous use.
  • Keep source sets clean. Superseded documents left in the index are a reliable way to get confidently outdated answers.
  • Notice vagueness. When a RAG tool goes general and abstract, retrieval probably came back empty. That is your cue to rephrase.

Where you have already met it

RAG is behind more tools than most people realize: chat interfaces that answer questions about an uploaded PDF, customer support bots that quote a company's real documentation, NotebookLM restricting itself to sources you provide, and AI search tools like Perplexity that retrieve pages before summarizing them.

Recognizing the pattern is useful, because it tells you what kind of mistakes to expect. Any tool that searches first and generates second will fail at the retrieval step sometimes — and knowing that is what turns a citation from reassurance into something you actually check.

See also What Is Fine-Tuning? for the other main way models are adapted, and How to Fact-Check AI Output.