Modern language models can accept very long inputs, sometimes hundreds of thousands of tokens or more. Does that make retrieval unnecessary?
When Long Context Works Well
- A small, fixed set of documents that fits comfortably.
- Tasks needing the whole document: summarising a contract, comparing sections, finding inconsistencies.
- Prototypes, where simplicity matters.
When RAG Is Still Better
- Large or growing collections that can't fit in any context window.
- Cost: every request pays for every token in the context; retrieval sends only relevant passages.
- Latency: long inputs take longer to process.
- Accuracy: models can overlook information buried in very long contexts; focused context often gives better answers.
- Permissions: retrieval can filter content per user.
- Freshness: indexes update incrementally.
Prompt Caching Changes the Maths
Where providers support caching repeated prompt prefixes, placing a stable document set at the start of the context can make long-context approaches cheaper for repeated questions.
Combining Them
Retrieve generously — larger chunks or whole sections — and rely on the long context to hold them. This reduces the risk of retrieval missing context.
Decide by Measuring
Compare accuracy, cost and latency on your own questions with both approaches. The answer depends on your document set and usage.