Fast retrieval methods — keyword search and vector search — are good at finding plausible candidates but imperfect at ordering them. A re-ranker fixes the order.
Two-Stage Retrieval
- Retrieve a broad set of candidates quickly, say the top 50–100.
- Re-rank those candidates with a slower, more accurate model, and keep the best few.
How Re-Rankers Work
Embedding search compares a query vector with document vectors computed separately. A cross-encoder re-ranker reads the query and each candidate together, so it can judge relevance much more precisely. Language models can also be used as re-rankers by asking them to score or order passages.
Benefits
- Better top results, which matters most when only a few passages go into an LLM prompt.
- Recovers relevant documents that the first stage ranked too low.
- Lets you retrieve more broadly without flooding the prompt.
Costs
Re-ranking adds latency and compute proportional to the number of candidates. Keep the candidate set moderate and measure the end-to-end effect.
Where It Helps Most
Question answering over large or varied document collections, customer support search, and RAG systems where irrelevant context leads to poor answers.
Measure It
Compare metrics such as recall@k and the rate of correct final answers with and without re-ranking on a realistic test set.