Conversion Rate Optimization (CRO)A/B testing statistics and experimentation platforms · Lesson 15 of 20
Sample size, significance and power
Video lecture
Sample size, significance and power
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Sample size, significance and power
Flip a fair coin a hundred times. Will you get exactly fifty heads? Probably not. You might get forty-six, or fifty-five. Now imagine two identical web pages. Their conversion rates will differ too, purely by chance. So when your variant beats the control, how do you know it's real? That's what this lecture is about. You'll learn the four inputs to every sample size calculation, what significance and power really mean, how to turn sample size into a test duration, and why confidence intervals beat a simple significant or not. Then you'll watch me build a sample size calculator in Python.
0:44 Why statistics matter
Why does this matter? Because without statistics, you'll ship noise. You'll celebrate random wins, then wonder why revenue didn't move. Or you'll give up on good ideas because a test was too small to detect them. Getting sample size right before you launch protects you from both. It also tells you, in advance, whether a test is even worth running on a given page, which saves months.
1:13 Four inputs
Every sample size calculation has four inputs. The baseline conversion rate: how often the control converts today. The minimum detectable effect: the smallest improvement worth detecting, often expressed as a relative lift, like ten percent. The significance level, alpha: the false positive risk you accept, usually five percent. And statistical power: the chance of detecting a real effect of that size, usually eighty percent. Change any one, and the sample you need changes.
1:45 Significance in plain words
Significance in plain language. A result is significant at five percent when a difference this large would happen less than five percent of the time if there were truly no difference. That's what the p-value is. A p-value of zero point zero three means: if the variants were identical, we'd see a difference at least this big about three percent of the time. What it does not mean: there's a ninety-five percent chance the variant is better. And it says nothing about how big or valuable the effect is.
2:24 Power in plain words
Power in plain language. Power is the probability that your test detects a real effect of at least the minimum detectable effect. With eighty percent power, if the true uplift is exactly your MDE, you'll detect it about four times out of five, and miss it one time in five. Underpowered tests end inconclusive even when the change genuinely helps. That's why a flat result from a tiny test tells you very little.
2:56 The numbers (5% alpha, 80% power)
Here are the numbers, from the full calculation at five percent significance and eighty percent power. They're illustrative planning figures. A three percent baseline and a ten percent relative lift needs about fifty-three thousand visitors per variant. Double the effect to twenty percent, and it drops to about fourteen thousand. A ten percent baseline with a ten percent lift needs about fifteen thousand. But a one percent baseline with a ten percent lift needs over a hundred and sixty thousand per variant. Two lessons. Halving the effect you want to detect roughly quadruples the sample. And low baseline rates need far more traffic.
3:41 From sample to duration
Turn sample size into duration. Divide the total sample, all variants together, by your daily eligible visitors. Then round up to whole weeks, so every day of the week is represented, because weekday and weekend behaviour differ. Avoid unusual periods like Eid, White Friday or Black Friday unless that's what you want to learn about. And if the duration comes out longer than about six to eight weeks, reconsider: test a bolder change, use a higher-frequency metric, or test on a larger audience. Long tests suffer from cookie loss and changing conditions.
4:21 Simple example: edtech course page (illustrative)
A simple example. A Pakistani edtech site's course page converts at three percent, with about four thousand eligible visitors a day. The team wants to detect a ten percent relative improvement. The calculator says about fifty-three thousand per variant, so a hundred and six thousand in total. At four thousand a day, that's twenty-seven days, rounded up to twenty-eight, four full weeks. That's feasible. If they'd wanted to detect a five percent improvement, it would take roughly four times as long. Too long. So they choose a bolder change worth a ten percent effect.
5:02 Realistic example: Riyadh electronics RPV (illustrative)
Now a realistic scenario about revenue. An electronics retailer in Riyadh wants revenue per visitor as the primary metric, which is usually better than conversion rate for a store. But revenue has much higher variance: most visitors spend nothing, and a few spend a lot. So the same relative effect needs more traffic. The analyst estimates the standard deviation from historical data, calculates the sample, and finds the test would take ten weeks. Options: test a bolder change, cap extreme order values as a pre-registered rule, or use a variance-reduction method like CUPED, which you'll meet in the next lesson but one. They choose CUPED plus capping, decided before launch.
5:50 Watch me do it: the calculator
Watch me build the calculator. I import square root and ceiling from the math module, and NormalDist from statistics, so there's nothing to install. The function takes the baseline, the relative MDE, alpha and power. It validates the inputs, because a baseline of three instead of zero point zero three would give nonsense. It computes the two conversion rates, the z-values for significance and power, and applies the standard two-proportion formula. A second function turns the sample into days, rounded up to whole weeks. I run it for four scenarios, and the printed table matches the numbers you saw earlier. Now I can answer is this test feasible? in seconds, in any planning meeting.
6:40 Confidence intervals
Report confidence intervals, not just significant or not. For example: an uplift of six percent, with a ninety-five percent interval from plus one to plus eleven percent. That tells people both the likely size and the uncertainty. If the interval includes zero, you can't rule out no effect. And remember practical significance. With huge traffic, tiny effects become significant but may not be worth the maintenance. Decide in advance what effect is worth acting on. That's your MDE.
7:14 Common mistakes
Common mistakes. Launching without a sample size calculation. Picking an MDE smaller than any realistic effect, creating impossibly long tests. Reading p equals zero point zero five as a ninety-five percent chance that B is better. Ignoring confidence intervals. Stopping on a calendar date regardless of sample. And using conversion rate sample sizes for revenue metrics.
7:38 Recap
Recap. Chance makes identical pages differ, so plan tests with four inputs: baseline, minimum detectable effect, significance and power. Halving the effect quadruples the sample, and low baselines need more traffic. Convert sample size into whole weeks, and rethink anything longer than six to eight weeks. Report effects with confidence intervals and decide in advance what's worth acting on. Try this now: copy the calculator from the lesson text, plug in your own page's baseline and traffic, and write down the smallest effect you could realistically detect in four weeks.
Why statistics matter
Conversion rates fluctuate naturally. If you flip a fair coin 100 times you will rarely get exactly 50 heads. Similarly, two identical pages will show different conversion rates purely by chance. Statistics help you decide whether an observed difference is likely to be real or just noise — and how much data you need to tell the difference.
Four inputs to every sample size calculation
| Input | Plain-language meaning | Common choice |
|---|---|---|
| Baseline conversion rate | How often the control converts today | From analytics |
| Minimum detectable effect (MDE) | The smallest improvement worth detecting | e.g. 10% relative uplift |
| Significance level (alpha) | Risk of a false positive you accept | 5% (95% confidence) |
| Statistical power | Chance of detecting a real effect of MDE size | 80% |
Significance in plain language
A result is "statistically significant at 5%" when a difference this large (or larger) would be unlikely — less than 5% probability — if there were truly no difference. It does not mean there is a 95% chance the variant is better, and it says nothing about how big or valuable the effect is.
The p-value is that probability. A p-value of 0.03 means: if the variants were truly identical, we would see a difference at least this large about 3% of the time.
Power in plain language
Power is the probability that your test detects a real effect of at least the MDE. With 80% power, if the true uplift is exactly your MDE, you will detect it about 4 times out of 5 — and miss it 1 time in 5. Underpowered tests often end "inconclusive" even when the change genuinely helps.
A rule of thumb for sample size
For a two-variant test at 5% significance and 80% power, a widely used approximation is:
n per variant ~ 16 x p x (1 - p) / d^2
p = baseline conversion rate (as a decimal)
d = absolute difference you want to detect (baseline x relative MDE)Worked example (illustrative):
Baseline p = 0.03 (3%)
Relative MDE = 10% -> d = 0.003
p(1-p) = 0.03 x 0.97 = 0.0291
n ~ 16 x 0.0291 / 0.000009 ~ 51,700 visitors per variant (~103,000 total)Now change one input:
| Baseline | Relative MDE | Approx. visitors per variant |
|---|---|---|
| 3% | 10% | ~51,700 |
| 3% | 20% | ~12,900 |
| 10% | 10% | ~14,400 |
| 1% | 10% | ~158,400 |
Key insights: halving the MDE roughly quadruples the sample needed, and low baseline rates need far more traffic. Use a proper sample size calculator for real planning; the rule of thumb is for intuition.
From sample size to duration
Duration (days) = Total sample needed / Daily eligible visitorsThen round up to whole weeks so every day of the week is represented equally — weekday and weekend behaviour often differ. Avoid running across unusual periods (major sales, holidays like Eid or Black Friday) unless that is what you want to learn about, and if you must, keep the test running for full cycles on either side.
If the duration is longer than about six to eight weeks, reconsider: test a bolder change, use a higher-frequency metric, or test on a larger audience. Very long tests suffer from cookie deletion and changing conditions.
Confidence intervals: the better summary
Rather than only "significant or not", report a confidence interval: for example, "uplift of +6%, with a 95% interval from +1% to +11%". It communicates both the likely size and the uncertainty. If the interval includes zero, you cannot rule out no effect.
Practical versus statistical significance
With huge traffic, tiny differences become statistically significant but may not justify the maintenance cost. With low traffic, valuable effects may not reach significance. Decide in advance what size of effect would be worth acting on — that is your MDE.
Hands-on: a sample size and duration calculator in Python
The rule of thumb above is for intuition. For planning, use the standard two-proportion formula. This calculator needs only the Python standard library:
from math import ceil, sqrt
from statistics import NormalDist
def sample_size_per_variant(baseline: float, relative_mde: float,
alpha: float = 0.05, power: float = 0.80) -> int:
"""Visitors per variant for a two-sided, two-proportion z-test (fixed horizon)."""
if not (0 < baseline < 1):
raise ValueError("baseline must be between 0 and 1, e.g. 0.03 for 3%")
p1, p2 = baseline, baseline * (1 + relative_mde)
if not (0 < p2 < 1):
raise ValueError("baseline x (1 + MDE) must stay between 0 and 1")
z_alpha = NormalDist().inv_cdf(1 - alpha / 2)
z_beta = NormalDist().inv_cdf(power)
p_bar = (p1 + p2) / 2
num = (z_alpha * sqrt(2 * p_bar * (1 - p_bar)) + z_beta * sqrt(p1 * (1 - p1) + p2 * (1 - p2))) ** 2
return ceil(num / (p2 - p1) ** 2)
def duration_days(n_per_variant: int, variants: int, daily_eligible_visitors: int) -> int:
days = ceil(n_per_variant * variants / daily_eligible_visitors)
return ceil(days / 7) * 7 # round up to whole weeks
if __name__ == "__main__":
for base, mde in [(0.03, 0.10), (0.03, 0.20), (0.10, 0.10), (0.01, 0.10)]:
n = sample_size_per_variant(base, mde)
print(f"baseline {base:.0%}, MDE {mde:.0%} -> {n:,} per variant; "
f"{duration_days(n, 2, 4_000)} days at 4,000 eligible visitors/day")Output:
baseline 3%, MDE 10% -> 53,211 per variant; 28 days at 4,000 eligible visitors/day
baseline 3%, MDE 20% -> 13,914 per variant; 7 days at 4,000 eligible visitors/day
baseline 10%, MDE 10% -> 14,751 per variant; 14 days at 4,000 eligible visitors/day
baseline 1%, MDE 10% -> 163,095 per variant; 84 days at 4,000 eligible visitors/dayThese are slightly higher than the 16·p(1−p)/d² approximation because the full formula accounts for the variant's own variance. Raising power to 90% (3%, 10% MDE) increases the requirement to about 71,000 per variant. Your platform's calculator may differ a little depending on its statistical method (for example sequential or Bayesian engines) — use the one that matches how you will analyse.
Revenue metrics need a different calculation
Revenue per visitor is usually a better primary metric for ecommerce than conversion rate, but its variance is much higher (most visitors spend nothing; a few spend a lot). Sample size for continuous metrics depends on the metric's standard deviation, which you estimate from historical data. Expect revenue metrics to need more traffic than conversion rate for the same relative MDE — and consider variance-reduction methods such as CUPED (lesson 5.4) or capping extreme values (winsorising) as pre-registered choices.
Common mistakes
- Launching without calculating sample size.
- Choosing an MDE smaller than any realistic effect, creating impossibly long tests.
- Interpreting p = 0.05 as "95% chance B is better".
- Ignoring confidence intervals and effect size.
- Stopping at a fixed calendar date regardless of sample reached.
Key takeaways
- Sample size depends on baseline rate, MDE, significance and power.
- Halving the MDE roughly quadruples the sample needed; low baselines need more traffic.
- Significance is about surprise under 'no difference', not the probability B is better.
- Run whole weeks and report confidence intervals, not just significant or not.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Use the rule of thumb (and then a sample size calculator) to estimate the duration of a test on one of your pages at 10% and 20% relative MDE.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.