Chunk size is one of the most influential settings in a RAG system, and there's no universal right answer.
The Trade-off
- Small chunks are precise: the retrieved text is focused. But they may lack context — a sentence saying "this limit applies" is useless without the limit.
- Large chunks carry context but blur the embedding and fill the prompt with irrelevant text.
Overlap
Overlapping consecutive chunks by a fraction (for example 10–20%) helps when an answer spans a boundary. Too much overlap wastes storage and returns near-duplicate results.
Structure Beats Fixed Sizes
Splitting at headings, paragraphs and list boundaries usually outperforms splitting every N characters. Use a size limit as a ceiling, not the main rule.
Test Systematically
- Build a set of realistic questions with the passages that answer them.
- Index the same documents with several chunking settings.
- Measure recall@k for each.
- Check end-to-end answer quality for the best candidates.
Different Content, Different Settings
FAQs suit one chunk per question. Long policies suit section-sized chunks. Code suits function-level chunks. It's fine to use different strategies for different sources.
Parent–Child Retrieval
Retrieve small, precise chunks, but pass their larger parent section to the model. This combines precise matching with enough context to answer.