Skip to content

Long Context Versus RAG

With context windows now very large, when should you just paste in the documents, and when is retrieval still better?

Editorial team 2 min read

Modern language models can accept very long inputs, sometimes hundreds of thousands of tokens or more. Does that make retrieval unnecessary?

When Long Context Works Well

  • A small, fixed set of documents that fits comfortably.
  • Tasks needing the whole document: summarising a contract, comparing sections, finding inconsistencies.
  • Prototypes, where simplicity matters.

When RAG Is Still Better

  • Large or growing collections that can't fit in any context window.
  • Cost: every request pays for every token in the context; retrieval sends only relevant passages.
  • Latency: long inputs take longer to process.
  • Accuracy: models can overlook information buried in very long contexts; focused context often gives better answers.
  • Permissions: retrieval can filter content per user.
  • Freshness: indexes update incrementally.

Prompt Caching Changes the Maths

Where providers support caching repeated prompt prefixes, placing a stable document set at the start of the context can make long-context approaches cheaper for repeated questions.

Combining Them

Retrieve generously — larger chunks or whole sections — and rely on the long context to hold them. This reduces the risk of retrieval missing context.

Decide by Measuring

Compare accuracy, cost and latency on your own questions with both approaches. The answer depends on your document set and usage.

More in RAG

All RAG guides →
RAG Guide · 1 min

RAG Architecture: The Components End to End

A map of a complete retrieval-augmented generation system, from ingestion to answer, and what each component is responsible for.

RAG 1 min read 6 Dec 2025

RAG Guide · 2 min

Document Parsing for RAG

Turning PDFs, slides, HTML and scans into clean, structured text — the unglamorous step that decides RAG quality.

RAG 2 min read 5 Dec 2025

RAG Guide · 2 min

Chunk Size and Overlap Tuning

How to choose chunk size and overlap for retrieval by testing against real questions rather than guessing.

RAG 2 min read 4 Dec 2025

RAG Guide · 2 min

Hybrid Search for RAG

Combining keyword and vector search so RAG finds both exact terms and paraphrased meaning.

RAG 2 min read 3 Dec 2025