When users search a product catalogue, a help centre or a document library, the order of results matters enormously. Learning to rank uses machine learning to order results by relevance.
Signals Used for Ranking
- Text relevance: keyword scores such as BM25 and semantic similarity from embeddings.
- Item quality: popularity, ratings, freshness, completeness.
- User context: location, language, past behaviour.
- Business rules: availability, margins, promotions.
Approaches
- Pointwise: predict a relevance score for each item.
- Pairwise: learn which of two items should rank higher.
- Listwise: optimise the quality of the whole ranked list directly.
Gradient-boosted ranking models (such as LambdaMART) are a common, strong choice.
Training Data
Relevance judgements from experts, or implicit feedback such as clicks and purchases. Clicks are biased — users click what's shown at the top — so corrections for position bias matter.
Evaluation
- Offline: NDCG, MRR and recall@k on judged queries.
- Online: A/B tests measuring clicks, conversions, reformulated searches and zero-result rates.
Practical Tips
- Start with a strong text-relevance baseline.
- Monitor queries that return no results or poor results.
- Add synonyms and spelling correction before complex models.
- Watch for feedback loops that entrench popular items.