---
title: "Peeking, pitfalls and reading results honestly"
description: "The peeking problem Most testing dashboards update continuously. It is tempting to check daily and stop the test the moment it shows \"95% significance\"…"
url: https://optimizeall.com/learn/conversion-rate-optimization/peeking-and-pitfalls
updated: 2026-10-05
---

Conversion Rate Optimization (CRO) · A/B testing statistics and experimentation platforms · lesson 16 of 20 · 11 min

# Peeking, pitfalls and reading results honestly

## The peeking problem

Most testing dashboards update continuously. It is tempting to check daily and stop the test the moment it shows "95% significance". This is called **peeking**, and with classical (fixed-horizon) statistics it dramatically inflates false positives. Early in a test, results swing wildly; if you check often enough, random noise will eventually cross the significance line.

Rules for fixed-horizon tests:

- Calculate sample size and duration **before** launch.
- Do not stop early because results look good (or bad), unless a guardrail shows real harm or the experience is broken.
- Analyse at the planned end.

## Sequential and Bayesian approaches

Some modern testing platforms use **sequential testing** methods designed for continuous monitoring, or **Bayesian** statistics that report the probability that a variant is better given the data and prior assumptions. These can legitimately allow earlier decisions — *if* you use the method as the tool intends. Understand which approach your tool uses and follow its rules; mixing approaches (peeking at a fixed-horizon test) is where errors creep in.

## Other common pitfalls

| Pitfall | What happens | Prevention |
|---|---|---|
| Multiple comparisons | Checking many metrics or segments means some will look "significant" by chance | One primary metric; treat segment findings as hypotheses |
| Novelty effect | Returning users react to anything new, temporarily | Run full cycles; check new vs returning |
| Primacy effect | Returning users initially resist change | Same as above; allow adaptation time |
| Seasonality / events | Unusual traffic periods distort results | Avoid or cover full cycles |
| Interaction with other tests | Overlapping tests contaminate each other | Coordinate or use mutually exclusive groups |
| Winner's curse | Winning tests tend to overestimate true effects | Expect real impact to be smaller; validate over time |
| Changing the test mid-flight | Altering variants or allocation invalidates data | Restart if you must change |

## Segment analysis after the test

After a test ends, it is useful to look at segments (device, new vs returning, traffic source) to **generate new hypotheses**. But segment "wins" in an overall flat test are often noise. If mobile appears to win strongly, run a follow-up test targeted at mobile rather than rolling out based on the slice.

## Interpreting outcomes

- **Clear winner** (primary metric improved, guardrails fine): roll out, document the mechanism, and consider follow-up tests that push the same idea further.
- **Clear loser**: valuable learning — the mechanism may be wrong, or the implementation poor. Document why you think it lost.
- **Inconclusive**: the effect, if any, is smaller than your MDE. Decide on other grounds (brand, cost, strategy) or test a bolder version. Do not call it a "trend toward significance".

## Worked example: the too-good-to-be-true win

A marketplace test shows a 35% uplift after three days, marked "significant". The team is tempted to stop. The analyst notes the planned sample was two weeks and that the test started during a payday weekend. They continue. By the planned end, the uplift settles at around 4% with a confidence interval that includes zero — inconclusive. Stopping early would have shipped a change on the basis of noise and a later "why didn't revenue rise 35%?" conversation.

## Reporting results honestly

A good test report includes:

```
Hypothesis and mechanism
Dates, audience, allocation, sample per variant, SRM check result
Primary metric: effect estimate + confidence interval (or credible interval)
Guardrails: status
Secondary metrics and notable segments (flagged as exploratory)
Decision and rationale
Learnings and next hypotheses
Screenshots of control and variant
```

## Sequential testing: what it allows, and its caveats

Many platforms now use statistics designed for continuous monitoring. For example, Optimizely's Stats Engine uses sequential methods that allow valid results when you check at any time; GrowthBook and Statsig offer sequential testing options for frequentist analyses; VWO's SmartStats and PostHog offer Bayesian analyses (check each tool's current documentation). Used correctly, these let you stop early when an effect is large.

Caveats that still apply:

| Caveat | Why it matters |
|---|---|
| **Wider intervals / less power at a given sample** | Always-valid methods pay for the freedom to peek; if you run to the same sample, a fixed-horizon test is usually more powerful |
| **Early stops overestimate effects** | Tests stopped early on a big swing tend to exaggerate the true effect (winner's curse); expect less in production |
| **Business cycles still matter** | Stopping after three days can miss weekend behaviour; many teams keep a minimum runtime of one to two full weeks even with sequential stats |
| **Novelty effects** | Early lift from returning users can fade |
| **Settings must be fixed in advance** | Changing metrics, allocation or the method mid-test invalidates guarantees |
| **Bayesian is not "peek-proof" by default** | Stopping rules based on "probability to beat control" can still inflate errors if priors and thresholds are not chosen with care |

Rule of thumb: **use the method your platform implements, the way it is documented, with a pre-registered minimum runtime and decision rule.**

## Hands-on: simulate the peeking problem

Seeing the problem is more convincing than reading about it. This simulation runs A/A tests (no real difference) and "peeks" daily, stopping at the first p < 0.05:

```python
import random
from math import sqrt
from statistics import NormalDist

def p_value(conv_a, n_a, conv_b, n_b):
    p = (conv_a + conv_b) / (n_a + n_b)
    se = sqrt(p * (1 - p) * (1 / n_a + 1 / n_b))
    if se == 0:
        return 1.0
    z = (conv_b / n_b - conv_a / n_a) / se
    return 2 * (1 - NormalDist().cdf(abs(z)))

random.seed(1)
RATE, DAILY, DAYS, SIMS = 0.03, 1_000, 28, 300
false_wins_peeking = false_wins_fixed = 0
for _ in range(SIMS):
    ca = cb = na = nb = 0
    stopped = False
    for day in range(DAYS):
        for _ in range(DAILY):
            if random.random() < 0.5:
                na += 1; ca += random.random() < RATE
            else:
                nb += 1; cb += random.random() < RATE
        if not stopped and p_value(ca, na, cb, nb) < 0.05:
            stopped = True                     # a "peeker" would stop and declare a winner
    false_wins_peeking += stopped
    false_wins_fixed += p_value(ca, na, cb, nb) < 0.05
print(f"False positives, fixed horizon: {false_wins_fixed / SIMS:.0%}")
print(f"False positives, daily peeking: {false_wins_peeking / SIMS:.0%}")
```

With no real difference, the fixed-horizon analysis flags a "winner" roughly 5% of the time, as designed; stopping at the first significant daily peek does so far more often. In one run (seed 1) the output was 6% for the fixed horizon and 29% for daily peeking — exact numbers vary with the seed and settings.

## Common mistakes

- Stopping when the dashboard first shows significance.
- Celebrating segment wins from an overall flat test.
- Reporting the point estimate as guaranteed future impact.
- Not recording losing and inconclusive tests.

## A pre-analysis checklist

Before you open the results dashboard at the planned end, confirm:

1. The planned sample size or duration has been reached.
2. The sample ratio check has passed.
3. Tracking worked for all variants throughout (no outages or broken events).
4. No other tests or major campaigns contaminated the audience without being accounted for.
5. The primary metric and decision rule are those written in the test plan.

Only then look at the primary metric. This ritual protects you from the most common ways good teams fool themselves.

## Video lecture: Peeking, pitfalls and reading results honestly

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Peeking, pitfalls and honest results
2. Why peeking hurts
3. The coin-toss analogy
4. Sequential and Bayesian options
5. Sequential caveats
6. Other pitfalls
7. Novelty and primacy
8. Simple example: the too-good win
9. Realistic example: UK subscription pricing (illustrative)
10. Watch me do it: simulate peeking
11. Segments and outcomes
12. Honest reporting
13. Common mistakes
14. Recap

## Lecture transcript

### Peeking, pitfalls and honest results

It's day three of your test. The dashboard shows ninety-five percent significance and a twenty-eight percent lift. Your manager says, great, ship it. Should you? Almost certainly not, and this lecture explains why. You'll learn the peeking problem, how sequential and Bayesian methods change the rules, and the caveats that still apply. You'll also learn the other pitfalls that fool good teams, like novelty effects, segment fishing and the winner's curse, and how to report results honestly. Then you'll watch me simulate peeking so you can see the problem with your own eyes.

### Why peeking hurts

Why does this matter? Because early results swing wildly. With small samples, random noise produces big apparent differences, which shrink as more data arrives. If you check a classic, fixed-horizon test every day and stop the first time it crosses the significance line, you dramatically increase the chance of declaring a winner when nothing changed. You ship noise, report inflated wins, and then can't explain why revenue didn't follow.

### The coin-toss analogy

Think of it like a coin-tossing contest. If you decide in advance to toss a hundred times and then count, you get a fair result. But if you're allowed to stop whenever you're ahead, you'll often declare victory for a fair coin. That's peeking. The fixed-horizon rules are simple: calculate the sample size before launch, don't stop early because results look good or bad, unless a guardrail shows real harm or the experience is broken, and analyse at the planned end.

### Sequential and Bayesian options

Modern platforms offer alternatives. Sequential testing methods are designed for continuous monitoring. Optimizely's Stats Engine, for example, uses sequential methods that stay valid when you look at any time, and tools like GrowthBook and Statsig offer sequential options. Bayesian approaches, used by VWO and available in PostHog, report the probability that a variant is better given the data. Used as the tool intends, these can legitimately allow earlier decisions. The error comes from mixing methods, like peeking at a fixed-horizon test.

### Sequential caveats

But sequential testing isn't magic. Five caveats. One: always-valid methods pay for the freedom to peek, usually with wider intervals or less power at the same sample. Two: tests stopped early on a big swing tend to overestimate the true effect, so expect less in production. Three: business cycles still matter, so many teams keep a minimum runtime of one to two full weeks. Four: novelty effects can fade. And five: settings must be fixed in advance. Bayesian results aren't automatically peek-proof either; stopping rules still need care.

### Other pitfalls

Other pitfalls. Multiple comparisons: check twenty metrics or segments and one will look significant by chance, so keep one primary metric. Novelty and primacy effects: returning users react to change, positively or negatively, for a while. Seasonality and events. Interactions with other tests running on the same pages. The winner's curse: winning tests tend to overestimate true effects. And changing the test mid-flight, which invalidates the data. If you must change it, restart.

### Novelty and primacy

Let's take a closer look at two effects that trick experienced teams. The novelty effect: returning visitors notice something new and click on it out of curiosity, so the variant looks great in week one and fades by week three. The primacy effect is the opposite: regulars are used to the old design and initially struggle with the new one, so the variant looks worse at first and recovers later. Both are strongest for sites with many returning users, like marketplaces and subscription services. The defences are the same. Run full weekly cycles, compare new and returning visitors separately, and for important changes, keep watching after launch with a holdback group.

### Simple example: the too-good win

A simple example. A marketplace test shows a thirty-five percent uplift after three days, flagged significant on a fixed-horizon dashboard. The planned sample was two weeks, and the test started on a payday weekend. The analyst insists on continuing. By the planned end, the uplift settles around four percent, with a confidence interval that includes zero. Inconclusive. Stopping early would have shipped a change on the basis of noise, and started a very uncomfortable conversation about why revenue didn't rise thirty-five percent.

### Realistic example: UK subscription pricing (illustrative)

Now a realistic scenario with illustrative details. A UK subscription brand uses a platform with sequential statistics. A new pricing page variant crosses the stopping boundary after six days with a large lift. The team is allowed to stop, statistically. But their pre-registered rule says: minimum runtime of fourteen days to cover two weekends and a pay cycle. They keep it running. At fourteen days the variant still wins, but the effect is about half the early estimate. They ship it, and report the smaller, more reliable number. They also plan a holdback to check the effect over the next two months, because pricing effects can change as the customer mix shifts.

### Watch me do it: simulate peeking

Watch me simulate peeking. I write a short Python script that runs three hundred A/A tests, where both variants are identical, with a three percent conversion rate and a thousand visitors a day for twenty-eight days. Each simulated day, a peeker checks the p-value and stops at the first result below zero point zero five. A disciplined analyst only looks at the end. When I run it, the fixed-horizon approach flags a false winner roughly five percent of the time, exactly as designed. The daily peeker flags false winners far more often. Same data, different behaviour, very different error rates. That's the whole argument, in one printout.

### Segments and outcomes

Segment analysis after a test is useful for generating ideas, and dangerous for making decisions. If an overall flat test shows mobile winning strongly, that's often noise, because you've looked at many slices. Treat it as a hypothesis: run a follow-up test targeted at mobile. And interpret outcomes carefully. A clear winner: roll out, and document the mechanism. A clear loser: valuable learning, so write down why you think it lost. Inconclusive: the effect, if any, is smaller than your MDE. Decide on other grounds, or test a bolder version. Never call it a trend towards significance.

### Honest reporting

Report results honestly. A good report includes the hypothesis and mechanism; dates, audience, allocation and sample per variant; the SRM check; the primary metric effect with its confidence or credible interval; guardrail status; secondary metrics and segments, flagged as exploratory; the decision and why; learnings and next hypotheses; and screenshots of each variant. And before you open the dashboard at the end, confirm the sample was reached, SRM passed, tracking worked throughout, and the decision rule is the one you wrote down.

### Common mistakes

Common mistakes. Stopping when the dashboard first shows significance. Using a sequential tool but ignoring business cycles. Treating Bayesian probability to beat control as permission to stop whenever it looks good. Celebrating segment wins from flat tests. Reporting early estimates as guaranteed future impact. And not recording losing or inconclusive tests, which are half of what you learn.

### Recap

Recap. Peeking at fixed-horizon tests multiplies false positives. Sequential and Bayesian methods can allow earlier decisions, but only when used as designed, with a pre-registered minimum runtime, and with the knowledge that early stops overestimate effects. Watch for multiple comparisons, novelty, seasonality, interactions and the winner's curse. Report effects with intervals, and treat segments as hypotheses. Try this now: run the peeking simulation from the lesson text, then check which statistical method your testing platform uses and write your team's stopping rule on one page.

## Key takeaways

- Peeking at fixed-horizon tests and stopping early inflates false positives.
- Sequential and Bayesian tools can support earlier decisions only when used as designed.
- Treat segment findings from a flat test as hypotheses, not wins.
- Winners tend to overestimate true effects; report intervals and document all outcomes.

## Try it

Review a past test (yours or a published case study) and write a report using the template, including whether peeking or multiple comparisons could have affected the conclusion.

- [Previous: Sample size, significance and power](https://optimizeall.com/learn/conversion-rate-optimization/sample-size-and-significance)
- [Next: Experimentation platforms, CUPED and variance reduction](https://optimizeall.com/learn/conversion-rate-optimization/experimentation-platforms-and-variance-reduction)
- [All lessons of Conversion Rate Optimization (CRO)](https://optimizeall.com/learn/conversion-rate-optimization)
