Conversion Rate Optimization (CRO)A/B testing statistics and experimentation platforms · Lesson 16 of 20

Peeking, pitfalls and reading results honestly

Article · 11 min · 9 min lecture

Video lecture

Peeking, pitfalls and reading results honestly

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Peeking, pitfalls and honest results

  • Day three, '95% significant'
  • Sequential and Bayesian methods
  • Caveats that still apply
  • A simulation you can run

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The peeking problem

Most testing dashboards update continuously. It is tempting to check daily and stop the test the moment it shows "95% significance". This is called peeking, and with classical (fixed-horizon) statistics it dramatically inflates false positives. Early in a test, results swing wildly; if you check often enough, random noise will eventually cross the significance line.

Rules for fixed-horizon tests:

  • Calculate sample size and duration before launch.
  • Do not stop early because results look good (or bad), unless a guardrail shows real harm or the experience is broken.
  • Analyse at the planned end.

Sequential and Bayesian approaches

Some modern testing platforms use sequential testing methods designed for continuous monitoring, or Bayesian statistics that report the probability that a variant is better given the data and prior assumptions. These can legitimately allow earlier decisions — if you use the method as the tool intends. Understand which approach your tool uses and follow its rules; mixing approaches (peeking at a fixed-horizon test) is where errors creep in.

Other common pitfalls

PitfallWhat happensPrevention
Multiple comparisonsChecking many metrics or segments means some will look "significant" by chanceOne primary metric; treat segment findings as hypotheses
Novelty effectReturning users react to anything new, temporarilyRun full cycles; check new vs returning
Primacy effectReturning users initially resist changeSame as above; allow adaptation time
Seasonality / eventsUnusual traffic periods distort resultsAvoid or cover full cycles
Interaction with other testsOverlapping tests contaminate each otherCoordinate or use mutually exclusive groups
Winner's curseWinning tests tend to overestimate true effectsExpect real impact to be smaller; validate over time
Changing the test mid-flightAltering variants or allocation invalidates dataRestart if you must change

Segment analysis after the test

After a test ends, it is useful to look at segments (device, new vs returning, traffic source) to generate new hypotheses. But segment "wins" in an overall flat test are often noise. If mobile appears to win strongly, run a follow-up test targeted at mobile rather than rolling out based on the slice.

Interpreting outcomes

  • Clear winner (primary metric improved, guardrails fine): roll out, document the mechanism, and consider follow-up tests that push the same idea further.
  • Clear loser: valuable learning — the mechanism may be wrong, or the implementation poor. Document why you think it lost.
  • Inconclusive: the effect, if any, is smaller than your MDE. Decide on other grounds (brand, cost, strategy) or test a bolder version. Do not call it a "trend toward significance".

Worked example: the too-good-to-be-true win

A marketplace test shows a 35% uplift after three days, marked "significant". The team is tempted to stop. The analyst notes the planned sample was two weeks and that the test started during a payday weekend. They continue. By the planned end, the uplift settles at around 4% with a confidence interval that includes zero — inconclusive. Stopping early would have shipped a change on the basis of noise and a later "why didn't revenue rise 35%?" conversation.

Reporting results honestly

A good test report includes:

Hypothesis and mechanism
Dates, audience, allocation, sample per variant, SRM check result
Primary metric: effect estimate + confidence interval (or credible interval)
Guardrails: status
Secondary metrics and notable segments (flagged as exploratory)
Decision and rationale
Learnings and next hypotheses
Screenshots of control and variant

Sequential testing: what it allows, and its caveats

Many platforms now use statistics designed for continuous monitoring. For example, Optimizely's Stats Engine uses sequential methods that allow valid results when you check at any time; GrowthBook and Statsig offer sequential testing options for frequentist analyses; VWO's SmartStats and PostHog offer Bayesian analyses (check each tool's current documentation). Used correctly, these let you stop early when an effect is large.

Caveats that still apply:

CaveatWhy it matters
Wider intervals / less power at a given sampleAlways-valid methods pay for the freedom to peek; if you run to the same sample, a fixed-horizon test is usually more powerful
Early stops overestimate effectsTests stopped early on a big swing tend to exaggerate the true effect (winner's curse); expect less in production
Business cycles still matterStopping after three days can miss weekend behaviour; many teams keep a minimum runtime of one to two full weeks even with sequential stats
Novelty effectsEarly lift from returning users can fade
Settings must be fixed in advanceChanging metrics, allocation or the method mid-test invalidates guarantees
Bayesian is not "peek-proof" by defaultStopping rules based on "probability to beat control" can still inflate errors if priors and thresholds are not chosen with care

Rule of thumb: use the method your platform implements, the way it is documented, with a pre-registered minimum runtime and decision rule.

Hands-on: simulate the peeking problem

Seeing the problem is more convincing than reading about it. This simulation runs A/A tests (no real difference) and "peeks" daily, stopping at the first p < 0.05:

import random
from math import sqrt
from statistics import NormalDist

def p_value(conv_a, n_a, conv_b, n_b):
    p = (conv_a + conv_b) / (n_a + n_b)
    se = sqrt(p * (1 - p) * (1 / n_a + 1 / n_b))
    if se == 0:
        return 1.0
    z = (conv_b / n_b - conv_a / n_a) / se
    return 2 * (1 - NormalDist().cdf(abs(z)))

random.seed(1)
RATE, DAILY, DAYS, SIMS = 0.03, 1_000, 28, 300
false_wins_peeking = false_wins_fixed = 0
for _ in range(SIMS):
    ca = cb = na = nb = 0
    stopped = False
    for day in range(DAYS):
        for _ in range(DAILY):
            if random.random() < 0.5:
                na += 1; ca += random.random() < RATE
            else:
                nb += 1; cb += random.random() < RATE
        if not stopped and p_value(ca, na, cb, nb) < 0.05:
            stopped = True                     # a "peeker" would stop and declare a winner
    false_wins_peeking += stopped
    false_wins_fixed += p_value(ca, na, cb, nb) < 0.05
print(f"False positives, fixed horizon: {false_wins_fixed / SIMS:.0%}")
print(f"False positives, daily peeking: {false_wins_peeking / SIMS:.0%}")

With no real difference, the fixed-horizon analysis flags a "winner" roughly 5% of the time, as designed; stopping at the first significant daily peek does so far more often. In one run (seed 1) the output was 6% for the fixed horizon and 29% for daily peeking — exact numbers vary with the seed and settings.

Common mistakes

  • Stopping when the dashboard first shows significance.
  • Celebrating segment wins from an overall flat test.
  • Reporting the point estimate as guaranteed future impact.
  • Not recording losing and inconclusive tests.

A pre-analysis checklist

Before you open the results dashboard at the planned end, confirm:

  1. The planned sample size or duration has been reached.
  2. The sample ratio check has passed.
  3. Tracking worked for all variants throughout (no outages or broken events).
  4. No other tests or major campaigns contaminated the audience without being accounted for.
  5. The primary metric and decision rule are those written in the test plan.

Only then look at the primary metric. This ritual protects you from the most common ways good teams fool themselves.

Key takeaways

  • Peeking at fixed-horizon tests and stopping early inflates false positives.
  • Sequential and Bayesian tools can support earlier decisions only when used as designed.
  • Treat segment findings from a flat test as hypotheses, not wins.
  • Winners tend to overestimate true effects; report intervals and document all outcomes.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A fixed-horizon test shows significance on day 3 of a planned 14. What should you do?
  2. An overall flat test shows a strong 'win' for tablet users. What is the best response?
  3. What is the winner's curse in A/B testing?
  4. Your platform uses sequential statistics and the variant crosses the stopping boundary on day 4. Which practice is still wise?

Put it into practice

Review a past test (yours or a published case study) and write a report using the template, including whether peeking or multiple comparisons could have affected the conclusion.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.