Skip to content

Regularisation: L1, L2 and Elastic Net

How penalising large coefficients reduces overfitting, and the practical difference between ridge, lasso and elastic net.

Editorial team 2 min read

Regularisation adds a penalty for model complexity to the training objective, discouraging extreme parameter values that fit noise.

L2 Regularisation (Ridge)

Penalises the sum of squared coefficients. It shrinks all coefficients towards zero smoothly but rarely makes any exactly zero. It is especially helpful when features are correlated, stabilising otherwise erratic coefficients.

L1 Regularisation (Lasso)

Penalises the sum of absolute coefficients. It tends to drive some coefficients exactly to zero, effectively performing feature selection and producing simpler models.

Elastic Net

Combines L1 and L2 penalties. It keeps lasso's ability to select features while behaving better when groups of correlated features exist.

The Strength Parameter

A hyperparameter — alpha in scikit-learn's Ridge and Lasso, or C (its inverse) in logistic regression and SVMs — controls how strong the penalty is. Too strong and the model underfits; too weak and it overfits. Choose it with cross-validation (RidgeCV, LassoCV).

Scale Features First

Penalties treat all coefficients equally, so features must be on comparable scales for regularisation to be fair.

In Neural Networks

The same ideas appear as weight decay (L2), alongside other regularisers such as dropout, early stopping and data augmentation.

Why It Works

By preferring smaller, simpler solutions, regularisation trades a little training accuracy for better performance on new data.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026