---
title: "Sample size, significance and power | Optimize All Academy"
description: "Why statistics matter Conversion rates fluctuate naturally. If you flip a fair coin 100 times you will rarely get exactly 50 heads. Similarly, two…"
url: https://optimizeall.com/learn/conversion-rate-optimization/sample-size-and-significance
updated: 2026-10-05
---

Conversion Rate Optimization (CRO) · A/B testing statistics and experimentation platforms · lesson 15 of 20 · 12 min

# Sample size, significance and power

## Why statistics matter

Conversion rates fluctuate naturally. If you flip a fair coin 100 times you will rarely get exactly 50 heads. Similarly, two identical pages will show different conversion rates purely by chance. Statistics help you decide whether an observed difference is likely to be **real** or just **noise** — and how much data you need to tell the difference.

## Four inputs to every sample size calculation

| Input | Plain-language meaning | Common choice |
|---|---|---|
| Baseline conversion rate | How often the control converts today | From analytics |
| Minimum detectable effect (MDE) | The smallest improvement worth detecting | e.g. 10% relative uplift |
| Significance level (alpha) | Risk of a false positive you accept | 5% (95% confidence) |
| Statistical power | Chance of detecting a real effect of MDE size | 80% |

## Significance in plain language

A result is "statistically significant at 5%" when a difference this large (or larger) would be unlikely — less than 5% probability — if there were truly no difference. It does **not** mean there is a 95% chance the variant is better, and it says nothing about how *big* or *valuable* the effect is.

The **p-value** is that probability. A p-value of 0.03 means: *if the variants were truly identical, we would see a difference at least this large about 3% of the time.*

## Power in plain language

**Power** is the probability that your test detects a real effect of at least the MDE. With 80% power, if the true uplift is exactly your MDE, you will detect it about 4 times out of 5 — and miss it 1 time in 5. Underpowered tests often end "inconclusive" even when the change genuinely helps.

## A rule of thumb for sample size

For a two-variant test at 5% significance and 80% power, a widely used approximation is:

```
n per variant  ~  16 x p x (1 - p) / d^2

p = baseline conversion rate (as a decimal)
d = absolute difference you want to detect (baseline x relative MDE)
```

Worked example (illustrative):

```
Baseline p = 0.03 (3%)
Relative MDE = 10%  ->  d = 0.003
p(1-p) = 0.03 x 0.97 = 0.0291
n ~ 16 x 0.0291 / 0.000009 ~ 51,700 visitors per variant (~103,000 total)
```

Now change one input:

| Baseline | Relative MDE | Approx. visitors per variant |
|---|---|---|
| 3% | 10% | ~51,700 |
| 3% | 20% | ~12,900 |
| 10% | 10% | ~14,400 |
| 1% | 10% | ~158,400 |

Key insights: **halving the MDE roughly quadruples the sample needed**, and **low baseline rates need far more traffic**. Use a proper sample size calculator for real planning; the rule of thumb is for intuition.

## From sample size to duration

```
Duration (days) = Total sample needed / Daily eligible visitors
```

Then round **up to whole weeks** so every day of the week is represented equally — weekday and weekend behaviour often differ. Avoid running across unusual periods (major sales, holidays like Eid or Black Friday) unless that is what you want to learn about, and if you must, keep the test running for full cycles on either side.

If the duration is longer than about six to eight weeks, reconsider: test a bolder change, use a higher-frequency metric, or test on a larger audience. Very long tests suffer from cookie deletion and changing conditions.

## Confidence intervals: the better summary

Rather than only "significant or not", report a **confidence interval**: for example, "uplift of +6%, with a 95% interval from +1% to +11%". It communicates both the likely size and the uncertainty. If the interval includes zero, you cannot rule out no effect.

## Practical versus statistical significance

With huge traffic, tiny differences become statistically significant but may not justify the maintenance cost. With low traffic, valuable effects may not reach significance. Decide in advance what size of effect would be **worth acting on** — that is your MDE.

## Hands-on: a sample size and duration calculator in Python

The rule of thumb above is for intuition. For planning, use the standard two-proportion formula. This calculator needs only the Python standard library:

```python
from math import ceil, sqrt
from statistics import NormalDist

def sample_size_per_variant(baseline: float, relative_mde: float,
                            alpha: float = 0.05, power: float = 0.80) -> int:
    """Visitors per variant for a two-sided, two-proportion z-test (fixed horizon)."""
    if not (0 < baseline < 1):
        raise ValueError("baseline must be between 0 and 1, e.g. 0.03 for 3%")
    p1, p2 = baseline, baseline * (1 + relative_mde)
    if not (0 < p2 < 1):
        raise ValueError("baseline x (1 + MDE) must stay between 0 and 1")
    z_alpha = NormalDist().inv_cdf(1 - alpha / 2)
    z_beta = NormalDist().inv_cdf(power)
    p_bar = (p1 + p2) / 2
    num = (z_alpha * sqrt(2 * p_bar * (1 - p_bar)) + z_beta * sqrt(p1 * (1 - p1) + p2 * (1 - p2))) ** 2
    return ceil(num / (p2 - p1) ** 2)

def duration_days(n_per_variant: int, variants: int, daily_eligible_visitors: int) -> int:
    days = ceil(n_per_variant * variants / daily_eligible_visitors)
    return ceil(days / 7) * 7                      # round up to whole weeks

if __name__ == "__main__":
    for base, mde in [(0.03, 0.10), (0.03, 0.20), (0.10, 0.10), (0.01, 0.10)]:
        n = sample_size_per_variant(base, mde)
        print(f"baseline {base:.0%}, MDE {mde:.0%} -> {n:,} per variant; "
              f"{duration_days(n, 2, 4_000)} days at 4,000 eligible visitors/day")
```

Output:

```
baseline 3%, MDE 10% -> 53,211 per variant; 28 days at 4,000 eligible visitors/day
baseline 3%, MDE 20% -> 13,914 per variant; 7 days at 4,000 eligible visitors/day
baseline 10%, MDE 10% -> 14,751 per variant; 14 days at 4,000 eligible visitors/day
baseline 1%, MDE 10% -> 163,095 per variant; 84 days at 4,000 eligible visitors/day
```

These are slightly higher than the 16·p(1−p)/d² approximation because the full formula accounts for the variant's own variance. Raising power to 90% (3%, 10% MDE) increases the requirement to about 71,000 per variant. Your platform's calculator may differ a little depending on its statistical method (for example sequential or Bayesian engines) — use the one that matches how you will analyse.

## Revenue metrics need a different calculation

Revenue per visitor is usually a better primary metric for ecommerce than conversion rate, but its variance is much higher (most visitors spend nothing; a few spend a lot). Sample size for continuous metrics depends on the metric's standard deviation, which you estimate from historical data. Expect revenue metrics to need **more** traffic than conversion rate for the same relative MDE — and consider variance-reduction methods such as CUPED (lesson 5.4) or capping extreme values (winsorising) as pre-registered choices.

## Common mistakes

- Launching without calculating sample size.
- Choosing an MDE smaller than any realistic effect, creating impossibly long tests.
- Interpreting p = 0.05 as "95% chance B is better".
- Ignoring confidence intervals and effect size.
- Stopping at a fixed calendar date regardless of sample reached.

## Video lecture: Sample size, significance and power

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. Sample size, significance and power
2. Why statistics matter
3. Four inputs
4. Significance in plain words
5. Power in plain words
6. The numbers (5% alpha, 80% power)
7. From sample to duration
8. Simple example: edtech course page (illustrative)
9. Realistic example: Riyadh electronics RPV (illustrative)
10. Watch me do it: the calculator
11. Confidence intervals
12. Common mistakes
13. Recap

## Lecture transcript

### Sample size, significance and power

Flip a fair coin a hundred times. Will you get exactly fifty heads? Probably not. You might get forty-six, or fifty-five. Now imagine two identical web pages. Their conversion rates will differ too, purely by chance. So when your variant beats the control, how do you know it's real? That's what this lecture is about. You'll learn the four inputs to every sample size calculation, what significance and power really mean, how to turn sample size into a test duration, and why confidence intervals beat a simple significant or not. Then you'll watch me build a sample size calculator in Python.

### Why statistics matter

Why does this matter? Because without statistics, you'll ship noise. You'll celebrate random wins, then wonder why revenue didn't move. Or you'll give up on good ideas because a test was too small to detect them. Getting sample size right before you launch protects you from both. It also tells you, in advance, whether a test is even worth running on a given page, which saves months.

### Four inputs

Every sample size calculation has four inputs. The baseline conversion rate: how often the control converts today. The minimum detectable effect: the smallest improvement worth detecting, often expressed as a relative lift, like ten percent. The significance level, alpha: the false positive risk you accept, usually five percent. And statistical power: the chance of detecting a real effect of that size, usually eighty percent. Change any one, and the sample you need changes.

### Significance in plain words

Significance in plain language. A result is significant at five percent when a difference this large would happen less than five percent of the time if there were truly no difference. That's what the p-value is. A p-value of zero point zero three means: if the variants were identical, we'd see a difference at least this big about three percent of the time. What it does not mean: there's a ninety-five percent chance the variant is better. And it says nothing about how big or valuable the effect is.

### Power in plain words

Power in plain language. Power is the probability that your test detects a real effect of at least the minimum detectable effect. With eighty percent power, if the true uplift is exactly your MDE, you'll detect it about four times out of five, and miss it one time in five. Underpowered tests end inconclusive even when the change genuinely helps. That's why a flat result from a tiny test tells you very little.

### The numbers (5% alpha, 80% power)

Here are the numbers, from the full calculation at five percent significance and eighty percent power. They're illustrative planning figures. A three percent baseline and a ten percent relative lift needs about fifty-three thousand visitors per variant. Double the effect to twenty percent, and it drops to about fourteen thousand. A ten percent baseline with a ten percent lift needs about fifteen thousand. But a one percent baseline with a ten percent lift needs over a hundred and sixty thousand per variant. Two lessons. Halving the effect you want to detect roughly quadruples the sample. And low baseline rates need far more traffic.

### From sample to duration

Turn sample size into duration. Divide the total sample, all variants together, by your daily eligible visitors. Then round up to whole weeks, so every day of the week is represented, because weekday and weekend behaviour differ. Avoid unusual periods like Eid, White Friday or Black Friday unless that's what you want to learn about. And if the duration comes out longer than about six to eight weeks, reconsider: test a bolder change, use a higher-frequency metric, or test on a larger audience. Long tests suffer from cookie loss and changing conditions.

### Simple example: edtech course page (illustrative)

A simple example. A Pakistani edtech site's course page converts at three percent, with about four thousand eligible visitors a day. The team wants to detect a ten percent relative improvement. The calculator says about fifty-three thousand per variant, so a hundred and six thousand in total. At four thousand a day, that's twenty-seven days, rounded up to twenty-eight, four full weeks. That's feasible. If they'd wanted to detect a five percent improvement, it would take roughly four times as long. Too long. So they choose a bolder change worth a ten percent effect.

### Realistic example: Riyadh electronics RPV (illustrative)

Now a realistic scenario about revenue. An electronics retailer in Riyadh wants revenue per visitor as the primary metric, which is usually better than conversion rate for a store. But revenue has much higher variance: most visitors spend nothing, and a few spend a lot. So the same relative effect needs more traffic. The analyst estimates the standard deviation from historical data, calculates the sample, and finds the test would take ten weeks. Options: test a bolder change, cap extreme order values as a pre-registered rule, or use a variance-reduction method like CUPED, which you'll meet in the next lesson but one. They choose CUPED plus capping, decided before launch.

### Watch me do it: the calculator

Watch me build the calculator. I import square root and ceiling from the math module, and NormalDist from statistics, so there's nothing to install. The function takes the baseline, the relative MDE, alpha and power. It validates the inputs, because a baseline of three instead of zero point zero three would give nonsense. It computes the two conversion rates, the z-values for significance and power, and applies the standard two-proportion formula. A second function turns the sample into days, rounded up to whole weeks. I run it for four scenarios, and the printed table matches the numbers you saw earlier. Now I can answer is this test feasible? in seconds, in any planning meeting.

### Confidence intervals

Report confidence intervals, not just significant or not. For example: an uplift of six percent, with a ninety-five percent interval from plus one to plus eleven percent. That tells people both the likely size and the uncertainty. If the interval includes zero, you can't rule out no effect. And remember practical significance. With huge traffic, tiny effects become significant but may not be worth the maintenance. Decide in advance what effect is worth acting on. That's your MDE.

### Common mistakes

Common mistakes. Launching without a sample size calculation. Picking an MDE smaller than any realistic effect, creating impossibly long tests. Reading p equals zero point zero five as a ninety-five percent chance that B is better. Ignoring confidence intervals. Stopping on a calendar date regardless of sample. And using conversion rate sample sizes for revenue metrics.

### Recap

Recap. Chance makes identical pages differ, so plan tests with four inputs: baseline, minimum detectable effect, significance and power. Halving the effect quadruples the sample, and low baselines need more traffic. Convert sample size into whole weeks, and rethink anything longer than six to eight weeks. Report effects with confidence intervals and decide in advance what's worth acting on. Try this now: copy the calculator from the lesson text, plug in your own page's baseline and traffic, and write down the smallest effect you could realistically detect in four weeks.

## Video transcript

Let's talk about how much data an A/B test actually needs, in plain language. Every test depends on four inputs. First, your baseline, meaning how often the page converts today. Second, the minimum detectable effect. That's the smallest improvement that would be worth knowing about. Third, the significance level, usually five percent. That's how much risk of a false alarm you're willing to accept. And fourth, power, usually eighty percent. That's your chance of catching a real improvement when there is one. Here's a handy rule of thumb. Visitors per variant come to roughly sixteen, times the baseline rate, times one minus the baseline rate, divided by the square of the difference you want to detect. So say your page converts at three percent and you want to detect a ten percent relative lift, from three percent to three point three. That works out to about fifty-two thousand visitors per variant. Now try detecting a twenty percent lift instead, and you only need around thirteen thousand. Halving the effect you're looking for roughly quadruples the traffic you need. That's why tiny tweaks on low-traffic pages almost never reach a clear answer. Once you know the sample size, divide it by your daily visitors to get the duration, then round up to whole weeks, because people behave differently on weekends. And when you report results, don't just say significant or not significant. Give the likely range, something like 'plus six percent, anywhere from plus one to plus eleven.' That range tells decision-makers how big the effect probably is, and how sure you really are.

## Key takeaways

- Sample size depends on baseline rate, MDE, significance and power.
- Halving the MDE roughly quadruples the sample needed; low baselines need more traffic.
- Significance is about surprise under 'no difference', not the probability B is better.
- Run whole weeks and report confidence intervals, not just significant or not.

## Try it

Use the rule of thumb (and then a sample size calculator) to estimate the duration of a test on one of your pages at 10% and 20% relative MDE.

- [Previous: A/B test fundamentals and experiment design](https://optimizeall.com/learn/conversion-rate-optimization/ab-test-fundamentals)
- [Next: Peeking, pitfalls and reading results honestly](https://optimizeall.com/learn/conversion-rate-optimization/peeking-and-pitfalls)
- [All lessons of Conversion Rate Optimization (CRO)](https://optimizeall.com/learn/conversion-rate-optimization)
