Users rarely phrase questions the way documents are written. Transforming the query before retrieval can improve recall considerably.
Conversational Context
In a chat, "What about for contractors?" means nothing on its own. Rewrite follow-up questions into standalone queries using the conversation history before searching.
Expansion
Add synonyms, acronyms and related terms: "PTO" → "paid time off, annual leave, vacation". A language model can generate variations, or you can maintain a domain glossary.
Multi-Query Retrieval
Generate several phrasings of the question, retrieve for each, and merge results. This catches documents that match one phrasing but not another.
Decomposition
Complex questions — "Compare the refund policies for annual and monthly plans" — can be split into sub-questions retrieved separately.
Hypothetical Document Embeddings (HyDE)
Ask a model to write a hypothetical answer, then embed that answer and search with it. The hypothetical answer often resembles real passages more closely than the question does. Watch that invented details don't mislead retrieval.
Costs
Each rewriting step adds latency and tokens. Use lightweight models for rewriting and apply techniques only where they improve measured recall.
Keep the Original
Always pass the user's original question to the generation step, so the answer addresses what was actually asked.