Regularisation adds a penalty for model complexity to the training objective, discouraging extreme parameter values that fit noise.
L2 Regularisation (Ridge)
Penalises the sum of squared coefficients. It shrinks all coefficients towards zero smoothly but rarely makes any exactly zero. It is especially helpful when features are correlated, stabilising otherwise erratic coefficients.
L1 Regularisation (Lasso)
Penalises the sum of absolute coefficients. It tends to drive some coefficients exactly to zero, effectively performing feature selection and producing simpler models.
Elastic Net
Combines L1 and L2 penalties. It keeps lasso's ability to select features while behaving better when groups of correlated features exist.
The Strength Parameter
A hyperparameter — alpha in scikit-learn's Ridge and Lasso, or C (its inverse) in logistic regression and SVMs — controls how strong the penalty is. Too strong and the model underfits; too weak and it overfits. Choose it with cross-validation (RidgeCV, LassoCV).
Scale Features First
Penalties treat all coefficients equally, so features must be on comparable scales for regularisation to be fair.
In Neural Networks
The same ideas appear as weight decay (L2), alongside other regularisers such as dropout, early stopping and data augmentation.
Why It Works
By preferring smaller, simpler solutions, regularisation trades a little training accuracy for better performance on new data.