Skip to content

AI for Data Analysis & Decision Making · Statistical traps, experiments and forecasting · lesson 13 of 16 · 8 min

Experiments and A/B tests: statistics you can defend

Why experiments beat clever analysis

Observational data tells you what happened together; a randomized experiment tells you what your change caused. Randomly assigning users (or stores, or regions) to control and treatment balances the hidden differences between groups, so a difference in outcomes can be attributed to the change, within a margin of chance. AI makes experiments easier to design and analyze, but it also makes it easier to run sloppy ones quickly. This lesson covers the statistics you need to read results honestly.

Five concepts in plain language

| Concept | Plain meaning | Common misreading | |---|---|---| | Baseline rate | Conversion (or other metric) in the control group | Using last year's rate instead of the test's own control | | Minimum detectable effect (MDE) | The smallest real lift the test is designed to detect | Setting it after seeing results | | Confidence interval | A range of lift values compatible with the data | Treating the point estimate as exact | | p-value | How surprising the data would be if there were truly no difference | "The probability the variant is better" (it is not) | | Statistical power | Chance of detecting the MDE if it is real (commonly 80%) | Ignoring it, then running underpowered tests |

Statistical significance is not business significance: a tiny, "significant" lift may not be worth the cost of change, and a large, non-significant lift from a small test may simply be unproven.

Hands-on: size the test before you start

# pip install statsmodels
from statsmodels.stats.power import NormalIndPower
from statsmodels.stats.proportion import proportion_effectsize

baseline = 0.030          # control conversion rate (from recent data)
mde_rel = 0.10            # want to detect a 10% relative lift (3.0% -> 3.3%)
target = baseline * (1 + mde_rel)

effect = proportion_effectsize(target, baseline)
n_per_group = NormalIndPower().solve_power(effect_size=effect, alpha=0.05, power=0.80, alternative="two-sided")
print(f"~{n_per_group:,.0f} visitors per group")

With these illustrative inputs the answer is in the tens of thousands of visitors per group. If your site cannot reach that in a reasonable time, test a bolder change (bigger MDE), use a higher-traffic page, or choose a different method.

Hands-on: read the result with an interval, not just a p-value

from statsmodels.stats.proportion import proportions_ztest, confint_proportions_2indep

conv = [690, 610]          # conversions: treatment, control (illustrative)
n    = [21000, 21000]      # visitors per group

stat, p = proportions_ztest(conv, n)
low, high = confint_proportions_2indep(conv[0], n[0], conv[1], n[1], compare="diff")
print(f"treatment {conv[0]/n[0]:.2%} vs control {conv[1]/n[1]:.2%}")
print(f"difference 95% CI: {low:.2%} to {high:.2%} (percentage points); p = {p:.3f}")

Report the interval: "The new checkout increased conversion by between roughly 0.05 and 0.7 percentage points (95% interval)". Then decide whether even the low end is worth it.

The rules that keep tests honest

  1. Decide before you start: hypothesis, primary metric, MDE, sample size, duration and decision rule.
  2. Randomize properly and check the split (sample ratio mismatch: a 50/50 test that arrives 52/48 on large numbers signals a bug).
  3. Don't peek and stop early when the p-value dips below 0.05; that inflates false positives. Run to the planned sample size, or use a method designed for sequential monitoring.
  4. Run full weeks to cover weekday patterns; avoid launching during unusual periods (a sale, Ramadan or Eid shifts) unless that is the context you want to learn about.
  5. One primary metric; treat secondary metrics and segment cuts as exploratory.
  6. Guardrail metrics (refund rate, page speed, complaints) so a "win" does not hide harm.

Worked example: a UK e-commerce checkout test

A London retailer tests a one-page checkout. Pre-registered plan: primary metric purchase conversion, MDE 20% relative (a bold change; about 15,000 visitors per group at this baseline), 80% power, two full weeks. After two weeks (illustrative, about 21,000 visitors per group), treatment converts at 3.29% versus 2.90%, 95% interval for the difference roughly 0.05 to 0.71 percentage points. Guardrails (refunds, support tickets) are flat. Even the low end pays for the change within a quarter, so they ship it, and they note that segment differences they saw afterward are hypotheses for the next test, not findings.

Second worked example: a Dubai SaaS pricing page

A SaaS company wants to test annual-plan framing but gets only a few thousand pricing-page visitors a month. The power calculation shows a small lift would need many months. Instead, they test a bolder change (a redesigned page with a clearer annual discount) with a larger MDE, and complement it with customer interviews. They avoid the trap of running an underpowered test and "learning" from noise.

Using AI in experiments

AI is useful for drafting the test plan, writing the power calculation and analysis code, checking for sample ratio mismatch, and drafting the readout with the interval and caveats. It should not choose the metric after the fact or slice results until something looks significant. Give it your pre-registered plan and ask it to analyze exactly that.

Pitfalls

  • Stopping as soon as p < 0.05.
  • Tiny tests on tiny traffic, then treating noise as learning.
  • Reporting only the point estimate.
  • Changing the primary metric after seeing results.

Video lecture: Experiments and A/B tests: statistics you can defend

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

  1. Experiments and A/B tests
  2. Why experiments
  3. Five concepts
  4. Power as a flashlight
  5. Significance ≠ importance
  6. Honest-test rules
  7. Timing
  8. Example 1: one-page checkout
  9. Example 2: low-traffic pricing page
  10. Watch me do it, part 1
  11. Watch me do it, part 2
  12. AI in experiments
  13. Common mistakes
  14. Healthy practice
  15. Recap and try this now

Lecture transcript

Experiments and A/B tests

Your team changed the checkout button color, and conversion went up. Did the color do it? Or was it the payday weekend, the new campaign, or random luck? Observational data can't tell you. A properly designed experiment can. In this lecture you'll learn why randomized experiments beat clever analysis, five statistics concepts in plain language, how to size a test before you start, how to read results with confidence intervals instead of just p-values, the rules that keep tests honest, and where AI helps.

Why experiments

Why do experiments matter so much? Because randomly assigning users, stores or regions to control and treatment balances the hidden differences between groups. So a difference in outcomes can be attributed to your change, within a margin of chance. Everything else, from regressions to before-and-after comparisons, involves arguing about what else might explain the result. AI makes experiments easier to design and analyze. It also makes it easier to run sloppy ones quickly. The statistics are what keep you honest.

Five concepts

Five concepts, in plain language. The baseline rate: the conversion in your control group. The minimum detectable effect: the smallest real lift the test is designed to detect. The confidence interval: a range of lift values compatible with the data. The p-value: how surprising your data would be if there were truly no difference. It is not the probability that the variant is better. And statistical power: the chance of detecting your minimum effect if it's real, commonly set at eighty percent.

Power as a flashlight

Here's an analogy for power. Imagine looking for a small coin on a beach with a flashlight. A weak flashlight, like a small sample, means you'll often miss the coin even when it's there, and you might mistake a shell for it. A strong flashlight, a large sample, finds the coin reliably. Power is the strength of your flashlight. If you run a test without enough power, a null result means very little, and a positive result is more likely to be a shell.

Significance ≠ importance

And one crucial distinction. Statistical significance isn't business significance. A tiny lift can be statistically significant with enough traffic, yet not be worth the cost of the change. And a large lift from a small test can be not significant, which means unproven, not zero. So always ask two questions: is the effect real, and is it big enough to matter? The confidence interval answers both at once, which is why we report it.

Honest-test rules

Now the rules that keep tests honest. Decide before you start: hypothesis, primary metric, minimum effect, sample size, duration and decision rule. Randomize properly, and check the split; a fifty-fifty test arriving fifty-two, forty-eight on large numbers signals a bug, called a sample ratio mismatch. Don't peek and stop early when the p-value dips below point zero five, because that inflates false positives. Run full weeks. Use one primary metric. And add guardrail metrics, like refunds or complaints, so a win doesn't hide harm.

Timing

Timing matters in our markets too. Avoid launching tests during unusual periods, like a big sale, or when Ramadan and Eid shift behavior, unless that's specifically the context you want to learn about. Seasonal events change who visits and how they buy, so a result from an unusual fortnight may not hold the rest of the year. And note that the dates of these events move each year, so plan your test calendar with them in mind.

Example 1: one-page checkout

First example, a business case with illustrative numbers. A London retailer tests a one-page checkout. The pre-registered plan: purchase conversion as the primary metric, a twenty percent relative minimum effect, because it's a bold change, eighty percent power, two full weeks, which needs roughly fifteen thousand visitors per group. After two weeks, treatment converts at three point two nine percent versus two point nine. The ninety-five percent interval for the difference runs from about point zero five to point seven one percentage points. Guardrails are flat. Even the low end pays back within a quarter, so they ship it.

Example 2: low-traffic pricing page

Second example, a different lesson. A Dubai software company wants to test annual plan framing but gets only a few thousand pricing page visitors a month. The power calculation shows detecting a small lift would take many months. So instead of running an underpowered test and learning from noise, they test a bolder change, a redesigned page with a clearer annual discount, with a larger minimum effect, and complement it with customer interviews. Knowing when not to run a small test is part of the skill.

Watch me do it, part 1

Watch me size a test. With statsmodels, I set the baseline conversion at three percent, from recent data, and a minimum detectable effect of ten percent relative, so from three to three point three percent. I convert that into an effect size, then solve for the sample size per group at five percent significance and eighty percent power. With these illustrative inputs, it's tens of thousands of visitors per group. If my site can't reach that in a few weeks, I test a bolder change, pick a higher-traffic page, or choose a different method.

Watch me do it, part 2

Then I read the result properly. Six hundred ninety conversions from twenty-one thousand visitors in treatment, six hundred ten from twenty-one thousand in control. The z-test gives a p-value, but I focus on the ninety-five percent confidence interval for the difference in proportions. It comes out entirely above zero, roughly five hundredths of a point to seven tenths of a point. So I write: the new checkout increased conversion by between about point zero five and point seven percentage points. Then I ask whether even the low end is worth it.

AI in experiments

Where does AI help? Drafting the test plan, writing the power calculation and analysis code, checking for sample ratio mismatch, and drafting the readout with the interval and caveats. Where must it not be used? Choosing the metric after the fact, or slicing results until something looks significant. Give it your pre-registered plan and ask it to analyze exactly that. And treat any segment differences found afterward as hypotheses for the next test.

Common mistakes

Common mistakes. Stopping as soon as the p-value dips below point zero five. Tiny tests on tiny traffic, with noise treated as learning. Reporting only the point estimate. Changing the primary metric after seeing results. Ignoring sample ratio mismatch. And declaring victory without checking guardrails.

Healthy practice

How do you know your experimentation practice is healthy? Every test has a written plan before launch. Sample sizes come from power calculations. Readouts lead with intervals and guardrails. And a growing log of past tests, including the ones that showed no effect, informs what you test next. Null results are learning too, as long as the test had enough power.

Recap and try this now

Recap. Randomized experiments attribute effects to your change. Know the five concepts: baseline, minimum detectable effect, confidence interval, p-value and power. Size tests before you start, read results with intervals, and separate statistical from business significance. Follow the honest-test rules, and use AI for plans, code and readouts, not for fishing. Try this now: pick one change you want to test, write the pre-registered plan, and run the power calculation from the lesson with your real baseline.

Key takeaways

  • Randomized experiments balance hidden differences, so outcome differences can be attributed to the change.
  • Size tests with a baseline, minimum detectable effect, significance level and power before you start.
  • Report confidence intervals; statistical significance is not the same as business significance.
  • Pre-register the plan, avoid peeking and stopping early, check sample ratio mismatch and guardrails.

Try it

Write a pre-registered plan for one real test (hypothesis, metric, MDE, sample size, duration, decision rule) and run the lesson's power calculation with your baseline.