A model is only as good as the metric used to judge it. Choose the metric before building models, based on what errors cost.
Classification Metrics
- Accuracy — share of correct predictions. Misleading with imbalanced classes: a fraud model that always says "not fraud" can be 99.9% accurate and useless.
- Precision — of the cases predicted positive, how many were right. Important when false alarms are costly.
- Recall — of the actual positives, how many were found. Important when misses are costly, such as disease screening.
- F1 score — the harmonic mean of precision and recall.
- ROC AUC — how well the model ranks positives above negatives across all thresholds.
- PR AUC — more informative than ROC AUC when positives are rare.
Regression Metrics
- MAE (mean absolute error) — average size of errors, in the target's units. Easy to explain.
- RMSE (root mean squared error) — penalises large errors more heavily.
- MAPE — percentage error; breaks down near zero.
- R² — share of variance explained; useful but can be misleading on its own.
Matching Metrics to Costs
Ask what happens when the model is wrong in each direction. If a missed fraud costs far more than a false alarm, optimise recall at an acceptable precision. Sometimes the best approach is to assign money to each error type and minimise expected cost directly.
Always Compare to a Baseline
Report every metric alongside a simple baseline — predicting the most common class or the average value — so the number has context.