Skip to content

Chunking Documents for Retrieval

How to split long documents so search finds the right passage — chunk size, overlap, structure and metadata.

Editorial team 2 min read

Retrieval works on passages, not whole documents. How you split documents into chunks often decides whether a RAG system works.

Why Chunk

An embedding of a 40-page manual is a blurry average of everything in it. Smaller chunks give precise matches and let you pass only relevant text to the language model.

Choosing a Size

  • Too small (a sentence): chunks lose the context needed to understand them.
  • Too large (many pages): embeddings blur and prompts fill with irrelevant text.
  • A few hundred words is a common starting point, and must fit the embedding model's input limit. Test sizes against real questions.

Split Along Structure

Split by headings, sections and paragraphs rather than a fixed character count. Keep tables and lists intact where possible.

Overlap

Overlapping neighbouring chunks slightly helps when an answer straddles a boundary.

Add Context and Metadata

  • Prefix each chunk with its document title and section heading.
  • Store source, URL, date, version and access permissions as metadata for filtering and citations.

Special Content

  • Tables: convert to text rows or keep as small, self-contained chunks.
  • Code: split by function or class.
  • Scanned PDFs: need OCR first; check quality.

Keep It Fresh

Re-index changed documents, delete chunks from removed ones, and record which version was indexed.

Test It

Write realistic questions with known answers and check whether the right chunk appears in the top results before involving a language model.

More in Generative AI

All Generative AI guides →
Generative AI Guide · 2 min

Prompt Engineering Fundamentals

The building blocks of a good prompt — context, task, constraints and format — with before-and-after examples.

Generative AI 2 min read 24 Jul 2026

Generative AI Guide · 2 min

Few-Shot Prompting With Examples

Showing a model a few examples of the input and output you want is often clearer than describing it. How to choose good examples.

Generative AI 2 min read 23 Jul 2026

Generative AI Guide · 2 min

Getting Structured Output From LLMs

How to get JSON and other machine-readable output reliably from a language model, and how to validate it.

Generative AI 2 min read 22 Jul 2026

Generative AI Guide · 2 min

Why Language Models Hallucinate

What hallucination is, why it happens, and practical ways to reduce and catch it.

Generative AI 2 min read 21 Jul 2026