Statistical power is the probability that a test detects an effect if one really exists.
Why It Matters
Underpowered experiments often miss real effects. Worse, the "significant" results they do find tend to exaggerate effect sizes.
What Determines Power
- Sample size: more data increases power.
- Effect size: larger effects are easier to detect.
- Variability: noisy metrics need more data.
- Significance level: stricter thresholds reduce power.
Calculating Sample Size
Before an experiment, decide:
- The minimum effect worth detecting.
- The significance level, commonly 5%.
- The desired power, commonly 80%.
Then use a calculator or statistical library to find the required sample size.
Practical Implications
- Small improvements need large samples.
- Low-traffic sites may not be able to detect small effects; test bigger changes instead.
- Don't stop tests early when results look good — it inflates false positives.
Reducing Variance
Techniques such as using pre-experiment data as a covariate can reduce required sample sizes.