Context windows have grown from a few thousand tokens to hundreds of thousands or more. This changes what's possible.
What It Enables
- Analysing entire contracts, codebases or books at once.
- Long conversations and agent sessions.
- Many examples in a single prompt.
- Comparing multiple documents directly.
Limitations
- Cost and latency: more tokens mean higher cost and slower responses.
- Attention to detail: performance can drop for information buried in very long contexts.
- Reasoning across the whole: finding a fact is easier than synthesising many scattered facts.
Making It Work
- Put long documents first and questions last.
- Label documents clearly.
- Ask the model to quote relevant passages before answering.
- Use prompt caching for repeated large contexts.
Long Context Versus Retrieval
- Long context: simpler, good for one-off analysis of a manageable document set.
- Retrieval: better for very large collections, frequent queries and cost control.
Many systems combine both: retrieve relevant material into a generous context.
Test
Evaluate on your own documents and questions.