Conversion Rate Optimization (CRO)A/B testing statistics and experimentation platforms · Lesson 16 of 20
Peeking, pitfalls and reading results honestly
Video lecture
Peeking, pitfalls and reading results honestly
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Peeking, pitfalls and honest results
It's day three of your test. The dashboard shows ninety-five percent significance and a twenty-eight percent lift. Your manager says, great, ship it. Should you? Almost certainly not, and this lecture explains why. You'll learn the peeking problem, how sequential and Bayesian methods change the rules, and the caveats that still apply. You'll also learn the other pitfalls that fool good teams, like novelty effects, segment fishing and the winner's curse, and how to report results honestly. Then you'll watch me simulate peeking so you can see the problem with your own eyes.
0:40 Why peeking hurts
Why does this matter? Because early results swing wildly. With small samples, random noise produces big apparent differences, which shrink as more data arrives. If you check a classic, fixed-horizon test every day and stop the first time it crosses the significance line, you dramatically increase the chance of declaring a winner when nothing changed. You ship noise, report inflated wins, and then can't explain why revenue didn't follow.
1:10 The coin-toss analogy
Think of it like a coin-tossing contest. If you decide in advance to toss a hundred times and then count, you get a fair result. But if you're allowed to stop whenever you're ahead, you'll often declare victory for a fair coin. That's peeking. The fixed-horizon rules are simple: calculate the sample size before launch, don't stop early because results look good or bad, unless a guardrail shows real harm or the experience is broken, and analyse at the planned end.
1:45 Sequential and Bayesian options
Modern platforms offer alternatives. Sequential testing methods are designed for continuous monitoring. Optimizely's Stats Engine, for example, uses sequential methods that stay valid when you look at any time, and tools like GrowthBook and Statsig offer sequential options. Bayesian approaches, used by VWO and available in PostHog, report the probability that a variant is better given the data. Used as the tool intends, these can legitimately allow earlier decisions. The error comes from mixing methods, like peeking at a fixed-horizon test.
2:20 Sequential caveats
But sequential testing isn't magic. Five caveats. One: always-valid methods pay for the freedom to peek, usually with wider intervals or less power at the same sample. Two: tests stopped early on a big swing tend to overestimate the true effect, so expect less in production. Three: business cycles still matter, so many teams keep a minimum runtime of one to two full weeks. Four: novelty effects can fade. And five: settings must be fixed in advance. Bayesian results aren't automatically peek-proof either; stopping rules still need care.
2:58 Other pitfalls
Other pitfalls. Multiple comparisons: check twenty metrics or segments and one will look significant by chance, so keep one primary metric. Novelty and primacy effects: returning users react to change, positively or negatively, for a while. Seasonality and events. Interactions with other tests running on the same pages. The winner's curse: winning tests tend to overestimate true effects. And changing the test mid-flight, which invalidates the data. If you must change it, restart.
3:30 Novelty and primacy
Let's take a closer look at two effects that trick experienced teams. The novelty effect: returning visitors notice something new and click on it out of curiosity, so the variant looks great in week one and fades by week three. The primacy effect is the opposite: regulars are used to the old design and initially struggle with the new one, so the variant looks worse at first and recovers later. Both are strongest for sites with many returning users, like marketplaces and subscription services. The defences are the same. Run full weekly cycles, compare new and returning visitors separately, and for important changes, keep watching after launch with a holdback group.
4:18 Simple example: the too-good win
A simple example. A marketplace test shows a thirty-five percent uplift after three days, flagged significant on a fixed-horizon dashboard. The planned sample was two weeks, and the test started on a payday weekend. The analyst insists on continuing. By the planned end, the uplift settles around four percent, with a confidence interval that includes zero. Inconclusive. Stopping early would have shipped a change on the basis of noise, and started a very uncomfortable conversation about why revenue didn't rise thirty-five percent.
4:54 Realistic example: UK subscription pricing (illustrative)
Now a realistic scenario with illustrative details. A UK subscription brand uses a platform with sequential statistics. A new pricing page variant crosses the stopping boundary after six days with a large lift. The team is allowed to stop, statistically. But their pre-registered rule says: minimum runtime of fourteen days to cover two weekends and a pay cycle. They keep it running. At fourteen days the variant still wins, but the effect is about half the early estimate. They ship it, and report the smaller, more reliable number. They also plan a holdback to check the effect over the next two months, because pricing effects can change as the customer mix shifts.
5:43 Watch me do it: simulate peeking
Watch me simulate peeking. I write a short Python script that runs three hundred A/A tests, where both variants are identical, with a three percent conversion rate and a thousand visitors a day for twenty-eight days. Each simulated day, a peeker checks the p-value and stops at the first result below zero point zero five. A disciplined analyst only looks at the end. When I run it, the fixed-horizon approach flags a false winner roughly five percent of the time, exactly as designed. The daily peeker flags false winners far more often. Same data, different behaviour, very different error rates. That's the whole argument, in one printout.
6:30 Segments and outcomes
Segment analysis after a test is useful for generating ideas, and dangerous for making decisions. If an overall flat test shows mobile winning strongly, that's often noise, because you've looked at many slices. Treat it as a hypothesis: run a follow-up test targeted at mobile. And interpret outcomes carefully. A clear winner: roll out, and document the mechanism. A clear loser: valuable learning, so write down why you think it lost. Inconclusive: the effect, if any, is smaller than your MDE. Decide on other grounds, or test a bolder version. Never call it a trend towards significance.
7:12 Honest reporting
Report results honestly. A good report includes the hypothesis and mechanism; dates, audience, allocation and sample per variant; the SRM check; the primary metric effect with its confidence or credible interval; guardrail status; secondary metrics and segments, flagged as exploratory; the decision and why; learnings and next hypotheses; and screenshots of each variant. And before you open the dashboard at the end, confirm the sample was reached, SRM passed, tracking worked throughout, and the decision rule is the one you wrote down.
7:48 Common mistakes
Common mistakes. Stopping when the dashboard first shows significance. Using a sequential tool but ignoring business cycles. Treating Bayesian probability to beat control as permission to stop whenever it looks good. Celebrating segment wins from flat tests. Reporting early estimates as guaranteed future impact. And not recording losing or inconclusive tests, which are half of what you learn.
8:13 Recap
Recap. Peeking at fixed-horizon tests multiplies false positives. Sequential and Bayesian methods can allow earlier decisions, but only when used as designed, with a pre-registered minimum runtime, and with the knowledge that early stops overestimate effects. Watch for multiple comparisons, novelty, seasonality, interactions and the winner's curse. Report effects with intervals, and treat segments as hypotheses. Try this now: run the peeking simulation from the lesson text, then check which statistical method your testing platform uses and write your team's stopping rule on one page.
The peeking problem
Most testing dashboards update continuously. It is tempting to check daily and stop the test the moment it shows "95% significance". This is called peeking, and with classical (fixed-horizon) statistics it dramatically inflates false positives. Early in a test, results swing wildly; if you check often enough, random noise will eventually cross the significance line.
Rules for fixed-horizon tests:
- Calculate sample size and duration before launch.
- Do not stop early because results look good (or bad), unless a guardrail shows real harm or the experience is broken.
- Analyse at the planned end.
Sequential and Bayesian approaches
Some modern testing platforms use sequential testing methods designed for continuous monitoring, or Bayesian statistics that report the probability that a variant is better given the data and prior assumptions. These can legitimately allow earlier decisions — if you use the method as the tool intends. Understand which approach your tool uses and follow its rules; mixing approaches (peeking at a fixed-horizon test) is where errors creep in.
Other common pitfalls
| Pitfall | What happens | Prevention |
|---|---|---|
| Multiple comparisons | Checking many metrics or segments means some will look "significant" by chance | One primary metric; treat segment findings as hypotheses |
| Novelty effect | Returning users react to anything new, temporarily | Run full cycles; check new vs returning |
| Primacy effect | Returning users initially resist change | Same as above; allow adaptation time |
| Seasonality / events | Unusual traffic periods distort results | Avoid or cover full cycles |
| Interaction with other tests | Overlapping tests contaminate each other | Coordinate or use mutually exclusive groups |
| Winner's curse | Winning tests tend to overestimate true effects | Expect real impact to be smaller; validate over time |
| Changing the test mid-flight | Altering variants or allocation invalidates data | Restart if you must change |
Segment analysis after the test
After a test ends, it is useful to look at segments (device, new vs returning, traffic source) to generate new hypotheses. But segment "wins" in an overall flat test are often noise. If mobile appears to win strongly, run a follow-up test targeted at mobile rather than rolling out based on the slice.
Interpreting outcomes
- Clear winner (primary metric improved, guardrails fine): roll out, document the mechanism, and consider follow-up tests that push the same idea further.
- Clear loser: valuable learning — the mechanism may be wrong, or the implementation poor. Document why you think it lost.
- Inconclusive: the effect, if any, is smaller than your MDE. Decide on other grounds (brand, cost, strategy) or test a bolder version. Do not call it a "trend toward significance".
Worked example: the too-good-to-be-true win
A marketplace test shows a 35% uplift after three days, marked "significant". The team is tempted to stop. The analyst notes the planned sample was two weeks and that the test started during a payday weekend. They continue. By the planned end, the uplift settles at around 4% with a confidence interval that includes zero — inconclusive. Stopping early would have shipped a change on the basis of noise and a later "why didn't revenue rise 35%?" conversation.
Reporting results honestly
A good test report includes:
Hypothesis and mechanism
Dates, audience, allocation, sample per variant, SRM check result
Primary metric: effect estimate + confidence interval (or credible interval)
Guardrails: status
Secondary metrics and notable segments (flagged as exploratory)
Decision and rationale
Learnings and next hypotheses
Screenshots of control and variantSequential testing: what it allows, and its caveats
Many platforms now use statistics designed for continuous monitoring. For example, Optimizely's Stats Engine uses sequential methods that allow valid results when you check at any time; GrowthBook and Statsig offer sequential testing options for frequentist analyses; VWO's SmartStats and PostHog offer Bayesian analyses (check each tool's current documentation). Used correctly, these let you stop early when an effect is large.
Caveats that still apply:
| Caveat | Why it matters |
|---|---|
| Wider intervals / less power at a given sample | Always-valid methods pay for the freedom to peek; if you run to the same sample, a fixed-horizon test is usually more powerful |
| Early stops overestimate effects | Tests stopped early on a big swing tend to exaggerate the true effect (winner's curse); expect less in production |
| Business cycles still matter | Stopping after three days can miss weekend behaviour; many teams keep a minimum runtime of one to two full weeks even with sequential stats |
| Novelty effects | Early lift from returning users can fade |
| Settings must be fixed in advance | Changing metrics, allocation or the method mid-test invalidates guarantees |
| Bayesian is not "peek-proof" by default | Stopping rules based on "probability to beat control" can still inflate errors if priors and thresholds are not chosen with care |
Rule of thumb: use the method your platform implements, the way it is documented, with a pre-registered minimum runtime and decision rule.
Hands-on: simulate the peeking problem
Seeing the problem is more convincing than reading about it. This simulation runs A/A tests (no real difference) and "peeks" daily, stopping at the first p < 0.05:
import random
from math import sqrt
from statistics import NormalDist
def p_value(conv_a, n_a, conv_b, n_b):
p = (conv_a + conv_b) / (n_a + n_b)
se = sqrt(p * (1 - p) * (1 / n_a + 1 / n_b))
if se == 0:
return 1.0
z = (conv_b / n_b - conv_a / n_a) / se
return 2 * (1 - NormalDist().cdf(abs(z)))
random.seed(1)
RATE, DAILY, DAYS, SIMS = 0.03, 1_000, 28, 300
false_wins_peeking = false_wins_fixed = 0
for _ in range(SIMS):
ca = cb = na = nb = 0
stopped = False
for day in range(DAYS):
for _ in range(DAILY):
if random.random() < 0.5:
na += 1; ca += random.random() < RATE
else:
nb += 1; cb += random.random() < RATE
if not stopped and p_value(ca, na, cb, nb) < 0.05:
stopped = True # a "peeker" would stop and declare a winner
false_wins_peeking += stopped
false_wins_fixed += p_value(ca, na, cb, nb) < 0.05
print(f"False positives, fixed horizon: {false_wins_fixed / SIMS:.0%}")
print(f"False positives, daily peeking: {false_wins_peeking / SIMS:.0%}")With no real difference, the fixed-horizon analysis flags a "winner" roughly 5% of the time, as designed; stopping at the first significant daily peek does so far more often. In one run (seed 1) the output was 6% for the fixed horizon and 29% for daily peeking — exact numbers vary with the seed and settings.
Common mistakes
- Stopping when the dashboard first shows significance.
- Celebrating segment wins from an overall flat test.
- Reporting the point estimate as guaranteed future impact.
- Not recording losing and inconclusive tests.
A pre-analysis checklist
Before you open the results dashboard at the planned end, confirm:
- The planned sample size or duration has been reached.
- The sample ratio check has passed.
- Tracking worked for all variants throughout (no outages or broken events).
- No other tests or major campaigns contaminated the audience without being accounted for.
- The primary metric and decision rule are those written in the test plan.
Only then look at the primary metric. This ritual protects you from the most common ways good teams fool themselves.
Key takeaways
- Peeking at fixed-horizon tests and stopping early inflates false positives.
- Sequential and Bayesian tools can support earlier decisions only when used as designed.
- Treat segment findings from a flat test as hypotheses, not wins.
- Winners tend to overestimate true effects; report intervals and document all outcomes.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Review a past test (yours or a published case study) and write a report using the template, including whether peeking or multiple comparisons could have affected the conclusion.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.