A distribution describes how likely different values are. Recognising common distributions helps you choose methods and spot anomalies.
Normal (Gaussian)
The symmetric bell curve. Heights, measurement errors and averages of many observations tend to be roughly normal. About 68% of values lie within one standard deviation of the mean and 95% within two.
Binomial
The number of successes in a fixed number of independent trials with the same probability — conversions out of 1,000 visitors.
Poisson
Counts of events in a fixed interval when events occur independently at a constant average rate — calls per hour, defects per batch. Its mean equals its variance.
Exponential
The time between independent events occurring at a constant rate — time between customer arrivals.
Uniform
Every value in a range is equally likely.
Long-Tailed Distributions
Log-normal and power-law distributions describe incomes, city sizes, file sizes and website traffic: most values are small, a few are enormous. Means are misleading here; use medians and percentiles, and consider log transforms.
Checking the Fit
Plot a histogram and, for normality, a Q–Q plot. Real data rarely matches a textbook distribution exactly; the question is whether it's close enough for your method.
Why It Matters
Many statistical tests assume particular distributions. Choosing methods that suit your data's shape avoids misleading results.