A loss function measures how wrong a model's predictions are. Training adjusts parameters to reduce it, so the choice shapes what the model learns.
Regression Losses
- Mean squared error: penalises large errors heavily.
- Mean absolute error: more robust to outliers.
- Huber loss: squared for small errors, absolute for large — a compromise.
- Quantile loss: for predicting percentiles and ranges.
Classification Losses
- Binary cross-entropy (log loss): for yes-or-no predictions.
- Categorical cross-entropy: for multi-class problems.
- Focal loss: focuses on hard examples, useful for imbalanced data.
- Hinge loss: used by support vector machines.
Other Losses
- Contrastive and triplet losses: for learning embeddings where similar items are close.
- Ranking losses: for search and recommendation.
Loss Versus Metric
The loss is what training optimises; the metric is what you care about. They should be aligned, but they're not always the same — you might train with cross-entropy but evaluate with F1.
Custom Losses
Weighting errors by business cost can align training with real outcomes.