Data breaks in many ways: a source system changes a field, a pipeline silently skips a day, a currency code appears that no one expected. Monitoring catches these before they mislead people.
What to Monitor
- Freshness: did data arrive on time?
- Volume: is the row count within its normal range?
- Schema: have columns been added, removed or retyped?
- Distribution: have averages, null rates or category mixes shifted unexpectedly?
- Business rules: totals reconcile, keys match, values are valid.
Rule-Based and Anomaly-Based Checks
Explicit rules catch known problems. Anomaly detection on metrics such as row counts and null rates catches unexpected ones. Use both.
Alerting Well
- Route alerts to the people who can fix the problem.
- Include context: which table, what changed, since when, which reports are affected.
- Tune thresholds to avoid alert fatigue.
- Distinguish warnings from critical failures.
Lineage
Knowing which reports and models depend on a table lets you assess impact quickly and notify affected users.
Ownership
Assign an owner to every important dataset, with agreed expectations (often called a data contract or service level) for freshness and quality.
Communicate Incidents
When data is wrong, tell users promptly, mark affected dashboards, and explain when it will be fixed.
Learn From Incidents
Fix root causes upstream and add checks so the same failure can't recur silently.