Skip to content

Feature Scaling and Normalisation

Why many algorithms need features on comparable scales, the difference between standardisation and min-max scaling, and when to skip it.

Editorial team 2 min read

Features often come in wildly different units — age in years, income in dollars, ratios between zero and one. Some algorithms care; others don't.

Which Algorithms Need Scaling

  • Distance-based: k-nearest neighbours, k-means, SVMs.
  • Gradient-based: linear and logistic regression with regularisation, neural networks.
  • PCA, which finds directions of maximum variance.

Tree-based models — decision trees, random forests, gradient boosting — split on thresholds and are unaffected by scaling.

Common Methods

  • Standardisation (z-score): subtract the mean and divide by the standard deviation. Features end up centred at zero with unit variance. A good default.
  • Min-max scaling: rescale to a fixed range, usually 0 to 1. Sensitive to outliers.
  • Robust scaling: use the median and interquartile range, reducing outlier influence.
  • Log transform: compress long-tailed features such as income or counts before scaling.

Avoid Leakage

Fit the scaler on the training data only, then apply the same transformation to validation and test data. Scaling the whole dataset first leaks information from test into training. Pipelines make this automatic.

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
model = make_pipeline(StandardScaler(), LogisticRegression()).fit(X_train, y_train)

Don't Forget Production

The fitted scaler is part of the model. Save it with the model so new data is transformed exactly as training data was.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026