Skip to content

Retrieval-Augmented Generation (RAG) Explained

How RAG grounds a language model's answers in your own documents, the components involved, and where it goes wrong.

Editorial team 2 min read

Retrieval-augmented generation (RAG) answers questions using your own documents: it retrieves relevant passages and gives them to a language model along with the question.

Why RAG

  • Models don't know your internal documents or recent information.
  • Grounding answers in retrieved text reduces hallucination and allows citations.
  • Updating knowledge means updating documents, not retraining a model.

The Components

  1. Ingestion: split documents into chunks and store them with metadata.
  2. Indexing: create embeddings for each chunk and store them in a vector index, often alongside a keyword index.
  3. Retrieval: embed the user's question and find the most relevant chunks.
  4. Generation: put those chunks in the prompt with instructions to answer from them and cite sources.

Where It Goes Wrong

Symptom Likely cause
Says the answer isn't available when it is Retrieval missed it: chunking, embeddings, or too few results
Confident answer not in the documents Weak grounding instructions or irrelevant chunks
Mixes up similar policies Chunks lack titles, dates or source metadata
Wrong after documents change Stale index

Improving It

  • Chunk along document structure and keep headings with each chunk.
  • Combine keyword and semantic search (hybrid search).
  • Re-rank retrieved chunks with a stronger model.
  • Filter by metadata and user permissions.

Evaluate the Two Halves

Measure retrieval (was the right chunk found?) separately from answer quality (was the answer correct and supported?). Fix retrieval first.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026