An A/B test randomly assigns users to a control (A) and a variant (B) to measure the effect of a change.
Before You Start
- State a hypothesis: "The shorter checkout will increase completed purchases."
- Choose a primary metric and a few guardrail metrics (refunds, page load time).
- Calculate the sample size needed to detect the smallest effect worth acting on.
- Fix the duration — typically whole weeks, to cover weekly cycles.
Randomise Properly
Assign at the right unit — usually the user, not the page view — so the same person always sees the same version. Check the split is as intended (a sample ratio mismatch signals a bug).
Common Mistakes
- Peeking: stopping as soon as results look significant inflates false positives. Decide the end in advance or use methods designed for continuous monitoring.
- Too many metrics: some will look significant by chance.
- Novelty effects: users react to anything new; longer tests help.
- Interference: users in different groups affecting each other, as in marketplaces or social features.
- Changing the test mid-way.
Analysing Results
Report the effect size with a confidence interval, check guardrail metrics, and look at important segments without over-interpreting small ones.
Making the Decision
Combine statistical results with cost, risk and strategy. A tiny, statistically significant gain may not be worth maintaining a complex feature.