Skip to content

k-Nearest Neighbours

A model that predicts by looking at the most similar examples it has seen. Simple, intuitive, and sensitive to scaling.

Editorial team 2 min read

k-nearest neighbours (k-NN) makes a prediction by finding the k training examples most similar to the new input and letting them decide.

How It Works

  • For classification, the neighbours vote and the most common class wins.
  • For regression, the prediction is the average of the neighbours' values.

There is no real training step: the model simply stores the data. All the work happens at prediction time.

Measuring Similarity

Similarity is usually Euclidean distance between feature vectors. Because distance depends on scale, features must be standardised first — otherwise a feature measured in thousands swamps one measured in fractions.

Choosing k

A small k (such as 1) follows the training data closely and overfits noise; a large k smooths predictions but can blur real boundaries. Choose k with cross-validation.

Strengths

  • Easy to understand and explain: "these five similar customers churned".
  • Naturally handles multi-class problems.
  • Works well when similar inputs really do have similar outputs.

Weaknesses

  • Slow predictions on large datasets (though approximate nearest-neighbour indexes help).
  • Suffers in high dimensions, where distances become less meaningful — the curse of dimensionality.
  • Sensitive to irrelevant features.

Where It Lives On

The same idea underpins vector search: finding the embeddings closest to a query is a nearest-neighbour search, the basis of semantic search and retrieval-augmented generation.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026