Retrieval-augmented generation (RAG) finds relevant material and supplies it to a model before it answers. Grounding ties an answer to evidence. Retrieval improves access to information; it does not guarantee a correct answer.
Follow the evidence through the pipeline
A typical document workflow splits material into passages, represents them for search, retrieves relevant passages and gives those passages to the model. Embeddings are numerical representations useful for similarity search. Keyword search and hybrid retrieval can also be useful; not every retrieval system needs a vector database.
The first failure can happen before generation: the right passage never reaches the model. The second can happen after retrieval: the model misreads a passage or adds a claim that it does not support. Measure those stages separately.
Try a tiny version without infrastructure
Write two short notes yourself. Select the relevant note manually and ask a chatbot to answer using only that note, quoting a short supporting passage. Ask an unrelated question next. The desired response is that the note does not contain the answer.
That exercise demonstrates evidence use, not a full automated RAG system. Once it works, explore indexing, retrieval quality, access controls and keeping documents current. Do not retrieve material a user is not permitted to access.
Sources and further reading
Read Google's embeddings documentation and our RAG guide.