Skip to content

Choosing the Right Evaluation Metric

Accuracy, precision, recall, F1, ROC AUC, MAE, RMSE: what each measures and how to choose one that matches the real cost of mistakes.

Editorial team 2 min read

A model is only as good as the metric used to judge it. Choose the metric before building models, based on what errors cost.

Classification Metrics

  • Accuracy — share of correct predictions. Misleading with imbalanced classes: a fraud model that always says "not fraud" can be 99.9% accurate and useless.
  • Precision — of the cases predicted positive, how many were right. Important when false alarms are costly.
  • Recall — of the actual positives, how many were found. Important when misses are costly, such as disease screening.
  • F1 score — the harmonic mean of precision and recall.
  • ROC AUC — how well the model ranks positives above negatives across all thresholds.
  • PR AUC — more informative than ROC AUC when positives are rare.

Regression Metrics

  • MAE (mean absolute error) — average size of errors, in the target's units. Easy to explain.
  • RMSE (root mean squared error) — penalises large errors more heavily.
  • MAPE — percentage error; breaks down near zero.
  • R² — share of variance explained; useful but can be misleading on its own.

Matching Metrics to Costs

Ask what happens when the model is wrong in each direction. If a missed fraud costs far more than a false alarm, optimise recall at an acceptable precision. Sometimes the best approach is to assign money to each error type and minimise expected cost directly.

Always Compare to a Baseline

Report every metric alongside a simple baseline — predicting the most common class or the average value — so the number has context.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026