Deployment is the beginning, not the end. Models degrade as data and behaviour change, often without obvious errors.
What to Monitor
- Operational health: latency, error rates, throughput, resource use and cost.
- Input data: missing values, new categories, out-of-range values and distribution shifts compared with training data.
- Predictions: distribution of scores and predicted classes over time.
- Outcomes: actual accuracy once true labels arrive — the most direct measure, though often delayed.
- Business metrics the model is meant to improve.
- Fairness: performance across important groups.
Delayed Labels
When outcomes arrive weeks later (did the loan default?), rely on input and prediction monitoring as early warnings, and evaluate accuracy as labels come in.
Alerting
Set thresholds for meaningful changes, route alerts to model owners with context, and avoid noisy alerts that get ignored.
Responding to Problems
- Check for pipeline or data-quality issues first — many "model problems" are broken inputs.
- Investigate genuine drift: what changed and why?
- Retrain, adjust thresholds, or fall back to a safer approach.
- Communicate with affected users.
Logging
Log inputs (with privacy safeguards), predictions, model versions and outcomes. Without logs you can't debug or audit decisions.
Review Regularly
Schedule periodic reviews of each production model: is it still accurate, fair, needed and owned?