Retrieval-augmented generation (RAG) is a way to make an LLM answer from your own documents. Before the model answers, the system searches a knowledge base for the most relevant passages and adds them to the prompt. The model then answers using that information, which makes answers more accurate, current and traceable.
Why do we need RAG?
An LLM only knows what it learned during training. It doesn't know your company's policies, last week's updates or a private document, and it may confidently make things up. RAG gives it the right information at the moment it answers.
How does RAG work, step by step?
- Split documents into chunks, such as paragraphs or sections.
- Create embeddings: turn each chunk into a vector that represents its meaning.
- Store them in a vector database such as FAISS, Chroma or Pinecone.
- Embed the user's question the same way.
- Retrieve the chunks most similar to the question.
- Generate: send the question and the retrieved chunks to the LLM with an instruction to answer only from them.
- Cite: show which sources the answer came from.
RAG vs fine-tuning
| RAG | Fine-tuning | |
|---|---|---|
| Best for | Adding knowledge (private or up-to-date) | Changing behaviour, tone or format |
| Updating | Just update the documents | Retrain the model |
| Shows sources | Yes | No |
| Cost to start | Lower | Higher |
A simple example
A college help desk bot could use RAG over the college's admission rules, fee notices and timetables. A student asks, "When is the last date to pay the second-semester fee?" The system retrieves the fee notice and the model answers from it, with a link to the notice.
Learn to build RAG
Program Zero teaches embeddings, vector databases and building a RAG pipeline in Phase 8, then deploys a RAG chatbot with CI/CD in Phase 10. For the background, read how LLMs work.