For classification, a single accuracy number hides what kinds of mistakes a model makes. The confusion matrix shows them.
The Confusion Matrix
For a binary problem:
| Predicted positive | Predicted negative | |
|---|---|---|
| Actually positive | True positive (TP) | False negative (FN) |
| Actually negative | False positive (FP) | True negative (TN) |
Precision and Recall
- Precision = TP / (TP + FP) — when the model says "positive", how often is it right?
- Recall = TP / (TP + FN) — of all real positives, how many did it catch?
The Trade-off
Most classifiers output a score or probability, and a threshold turns it into a decision. Lowering the threshold catches more positives (higher recall) but raises more false alarms (lower precision). Raising it does the opposite.
Choosing the Threshold
- Spam filter: favour precision — losing a real email is worse than seeing some spam.
- Cancer screening: favour recall — a missed case is far worse than an extra test.
- Fraud review: set the threshold so the number of flagged cases matches reviewer capacity.
A precision–recall curve shows every possible trade-off; pick the point that suits the decision.
Multi-Class Problems
For several classes, compute precision and recall per class and combine them with macro averaging (treat classes equally) or weighted averaging (by class size). Look at the full confusion matrix to see which classes are confused with each other.