k-nearest neighbours (k-NN) makes a prediction by finding the k training examples most similar to the new input and letting them decide.
How It Works
- For classification, the neighbours vote and the most common class wins.
- For regression, the prediction is the average of the neighbours' values.
There is no real training step: the model simply stores the data. All the work happens at prediction time.
Measuring Similarity
Similarity is usually Euclidean distance between feature vectors. Because distance depends on scale, features must be standardised first — otherwise a feature measured in thousands swamps one measured in fractions.
Choosing k
A small k (such as 1) follows the training data closely and overfits noise; a large k smooths predictions but can blur real boundaries. Choose k with cross-validation.
Strengths
- Easy to understand and explain: "these five similar customers churned".
- Naturally handles multi-class problems.
- Works well when similar inputs really do have similar outputs.
Weaknesses
- Slow predictions on large datasets (though approximate nearest-neighbour indexes help).
- Suffers in high dimensions, where distances become less meaningful — the curse of dimensionality.
- Sensitive to irrelevant features.
Where It Lives On
The same idea underpins vector search: finding the embeddings closest to a query is a nearest-neighbour search, the basis of semantic search and retrieval-augmented generation.