Anomaly detection finds observations that differ markedly from the norm. It's used for fraud detection, equipment monitoring, cybersecurity and data quality checks.
Types of Anomalies
- Point anomalies: a single unusual value, such as a huge transaction.
- Contextual anomalies: normal in one context but not another — high electricity use at 3am.
- Collective anomalies: a sequence that is unusual as a whole.
Approaches
- Statistical rules: z-scores, interquartile-range limits, control charts. Simple and explainable.
- Isolation Forest: isolates points with random splits; anomalies are isolated quickly.
- Local Outlier Factor: compares a point's density with its neighbours'.
- One-class SVM: learns a boundary around normal data.
- Autoencoders: neural networks that reconstruct normal data well and anomalies poorly.
- Supervised classifiers: when labelled anomalies exist, standard classification (with class imbalance handling) often works best.
The Challenge of Evaluation
Labelled anomalies are usually rare or missing. Use whatever confirmed cases exist, have experts review top-ranked alerts, and track precision on reviewed cases over time.
Practical Tips
- Define "normal" carefully; include seasonality and context.
- Tune the alert threshold to reviewers' capacity.
- Expect drift: what's normal changes, so retrain regularly.
- Combine anomaly scores with business rules to reduce false alarms.