Skip to content

Calibrating Predicted Probabilities

When a model says 80%, does it happen 80% of the time? How to check and fix probability calibration.

Editorial team 2 min read

Many classifiers output probabilities, and many decisions depend on them being accurate: pricing risk, prioritising reviews, combining predictions. Calibration measures whether those probabilities can be taken at face value.

What Calibration Means

A model is well calibrated if, among all the cases where it predicts 0.8, about 80% are actually positive. A model can rank cases well (high ROC AUC) and still be poorly calibrated.

How to Check

  • Reliability diagram: group predictions into bins (0–0.1, 0.1–0.2…) and plot the average prediction against the observed frequency in each bin. A calibrated model follows the diagonal.
  • Brier score: the mean squared difference between predicted probabilities and outcomes; lower is better.

Which Models Are Miscalibrated

  • Naive Bayes and boosted trees are often overconfident or skewed.
  • SVMs don't produce probabilities natively.
  • Class weighting and resampling for imbalance distort probabilities.
  • Logistic regression is usually reasonably calibrated.

How to Fix It

  • Platt scaling: fit a logistic regression on the model's scores.
  • Isotonic regression: a flexible, non-decreasing mapping; needs more data.

Fit the calibrator on data the model wasn't trained on, for example with CalibratedClassifierCV.

from sklearn.calibration import CalibratedClassifierCV
calibrated = CalibratedClassifierCV(model, method="isotonic", cv=5).fit(X_train, y_train)

When It Matters

Calibrate whenever probabilities feed a decision threshold, an expected-cost calculation, or a report to people who will read "70%" literally.

More in Machine learning

All Machine learning guides →
Machine learning Guide · 2 min

Linear Regression Explained

The simplest predictive model: how linear regression fits a line through data, how to read its coefficients, and when it breaks down.

Machine learning 2 min read 17 Sep 2026

Machine learning Guide · 2 min

Logistic Regression for Classification

Despite its name, logistic regression is a classification method. How it produces probabilities and why it remains a strong baseline.

Machine learning 2 min read 16 Sep 2026

Machine learning Guide · 2 min

Decision Trees

How decision trees split data with simple questions, why they are easy to explain, and why single trees overfit.

Machine learning 2 min read 15 Sep 2026

Machine learning Guide · 2 min

Random Forests

Why averaging many randomised decision trees produces a robust, accurate model with little tuning.

Machine learning 2 min read 14 Sep 2026