Conversion Rate Optimization (CRO)A/B testing statistics and experimentation platforms · Lesson 15 of 20

Sample size, significance and power

Video lesson · 12 min · 8 min lecture

Video lecture

Sample size, significance and power

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Sample size, significance and power

  • Chance makes identical pages differ
  • Four inputs to sample size
  • Significance and power in plain words
  • Watch: a Python calculator

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why statistics matter

Conversion rates fluctuate naturally. If you flip a fair coin 100 times you will rarely get exactly 50 heads. Similarly, two identical pages will show different conversion rates purely by chance. Statistics help you decide whether an observed difference is likely to be real or just noise — and how much data you need to tell the difference.

Four inputs to every sample size calculation

InputPlain-language meaningCommon choice
Baseline conversion rateHow often the control converts todayFrom analytics
Minimum detectable effect (MDE)The smallest improvement worth detectinge.g. 10% relative uplift
Significance level (alpha)Risk of a false positive you accept5% (95% confidence)
Statistical powerChance of detecting a real effect of MDE size80%

Significance in plain language

A result is "statistically significant at 5%" when a difference this large (or larger) would be unlikely — less than 5% probability — if there were truly no difference. It does not mean there is a 95% chance the variant is better, and it says nothing about how big or valuable the effect is.

The p-value is that probability. A p-value of 0.03 means: if the variants were truly identical, we would see a difference at least this large about 3% of the time.

Power in plain language

Power is the probability that your test detects a real effect of at least the MDE. With 80% power, if the true uplift is exactly your MDE, you will detect it about 4 times out of 5 — and miss it 1 time in 5. Underpowered tests often end "inconclusive" even when the change genuinely helps.

A rule of thumb for sample size

For a two-variant test at 5% significance and 80% power, a widely used approximation is:

n per variant  ~  16 x p x (1 - p) / d^2

p = baseline conversion rate (as a decimal)
d = absolute difference you want to detect (baseline x relative MDE)

Worked example (illustrative):

Baseline p = 0.03 (3%)
Relative MDE = 10%  ->  d = 0.003
p(1-p) = 0.03 x 0.97 = 0.0291
n ~ 16 x 0.0291 / 0.000009 ~ 51,700 visitors per variant (~103,000 total)

Now change one input:

BaselineRelative MDEApprox. visitors per variant
3%10%~51,700
3%20%~12,900
10%10%~14,400
1%10%~158,400

Key insights: halving the MDE roughly quadruples the sample needed, and low baseline rates need far more traffic. Use a proper sample size calculator for real planning; the rule of thumb is for intuition.

From sample size to duration

Duration (days) = Total sample needed / Daily eligible visitors

Then round up to whole weeks so every day of the week is represented equally — weekday and weekend behaviour often differ. Avoid running across unusual periods (major sales, holidays like Eid or Black Friday) unless that is what you want to learn about, and if you must, keep the test running for full cycles on either side.

If the duration is longer than about six to eight weeks, reconsider: test a bolder change, use a higher-frequency metric, or test on a larger audience. Very long tests suffer from cookie deletion and changing conditions.

Confidence intervals: the better summary

Rather than only "significant or not", report a confidence interval: for example, "uplift of +6%, with a 95% interval from +1% to +11%". It communicates both the likely size and the uncertainty. If the interval includes zero, you cannot rule out no effect.

Practical versus statistical significance

With huge traffic, tiny differences become statistically significant but may not justify the maintenance cost. With low traffic, valuable effects may not reach significance. Decide in advance what size of effect would be worth acting on — that is your MDE.

Hands-on: a sample size and duration calculator in Python

The rule of thumb above is for intuition. For planning, use the standard two-proportion formula. This calculator needs only the Python standard library:

from math import ceil, sqrt
from statistics import NormalDist

def sample_size_per_variant(baseline: float, relative_mde: float,
                            alpha: float = 0.05, power: float = 0.80) -> int:
    """Visitors per variant for a two-sided, two-proportion z-test (fixed horizon)."""
    if not (0 < baseline < 1):
        raise ValueError("baseline must be between 0 and 1, e.g. 0.03 for 3%")
    p1, p2 = baseline, baseline * (1 + relative_mde)
    if not (0 < p2 < 1):
        raise ValueError("baseline x (1 + MDE) must stay between 0 and 1")
    z_alpha = NormalDist().inv_cdf(1 - alpha / 2)
    z_beta = NormalDist().inv_cdf(power)
    p_bar = (p1 + p2) / 2
    num = (z_alpha * sqrt(2 * p_bar * (1 - p_bar)) + z_beta * sqrt(p1 * (1 - p1) + p2 * (1 - p2))) ** 2
    return ceil(num / (p2 - p1) ** 2)

def duration_days(n_per_variant: int, variants: int, daily_eligible_visitors: int) -> int:
    days = ceil(n_per_variant * variants / daily_eligible_visitors)
    return ceil(days / 7) * 7                      # round up to whole weeks

if __name__ == "__main__":
    for base, mde in [(0.03, 0.10), (0.03, 0.20), (0.10, 0.10), (0.01, 0.10)]:
        n = sample_size_per_variant(base, mde)
        print(f"baseline {base:.0%}, MDE {mde:.0%} -> {n:,} per variant; "
              f"{duration_days(n, 2, 4_000)} days at 4,000 eligible visitors/day")

Output:

baseline 3%, MDE 10% -> 53,211 per variant; 28 days at 4,000 eligible visitors/day
baseline 3%, MDE 20% -> 13,914 per variant; 7 days at 4,000 eligible visitors/day
baseline 10%, MDE 10% -> 14,751 per variant; 14 days at 4,000 eligible visitors/day
baseline 1%, MDE 10% -> 163,095 per variant; 84 days at 4,000 eligible visitors/day

These are slightly higher than the 16·p(1−p)/d² approximation because the full formula accounts for the variant's own variance. Raising power to 90% (3%, 10% MDE) increases the requirement to about 71,000 per variant. Your platform's calculator may differ a little depending on its statistical method (for example sequential or Bayesian engines) — use the one that matches how you will analyse.

Revenue metrics need a different calculation

Revenue per visitor is usually a better primary metric for ecommerce than conversion rate, but its variance is much higher (most visitors spend nothing; a few spend a lot). Sample size for continuous metrics depends on the metric's standard deviation, which you estimate from historical data. Expect revenue metrics to need more traffic than conversion rate for the same relative MDE — and consider variance-reduction methods such as CUPED (lesson 5.4) or capping extreme values (winsorising) as pre-registered choices.

Common mistakes

  • Launching without calculating sample size.
  • Choosing an MDE smaller than any realistic effect, creating impossibly long tests.
  • Interpreting p = 0.05 as "95% chance B is better".
  • Ignoring confidence intervals and effect size.
  • Stopping at a fixed calendar date regardless of sample reached.

Key takeaways

  • Sample size depends on baseline rate, MDE, significance and power.
  • Halving the MDE roughly quadruples the sample needed; low baselines need more traffic.
  • Significance is about surprise under 'no difference', not the probability B is better.
  • Run whole weeks and report confidence intervals, not just significant or not.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What happens to required sample size if you halve the minimum detectable effect?
  2. Which interpretation of p = 0.03 is correct?
  3. Why should tests run for whole weeks?
  4. Using the full two-proportion formula at 5% significance and 80% power, roughly how many visitors per variant are needed to detect a 10% relative lift on a 3% baseline?

Put it into practice

Use the rule of thumb (and then a sample size calculator) to estimate the duration of a test on one of your pages at 10% and 20% relative MDE.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.