Conversion Rate Optimization (CRO)A/B testing statistics and experimentation platforms · Lesson 17 of 20

Experimentation platforms, CUPED and variance reduction

Article · 8 min · 9 min lecture

Video lecture

Experimentation platforms, CUPED and variance reduction

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Experimentation platforms and CUPED

  • Life after Google Optimize
  • Four categories of tools
  • Choosing and validating a platform
  • Variance reduction with CUPED

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The experimentation stack after Google Optimize

Google Optimize and Optimize 360 were sunset on 30 September 2023. Google did not launch a direct replacement; instead GA4 integrates with third-party testing tools. The market now looks roughly like this (features change often — check vendors' current documentation and pricing):

CategoryExamplesTypical strengthsConsider
Marketer-led web experimentationVWO, Optimizely (Web Experimentation), AB Tasty, Kameleoon, ConvertVisual editors, targeting, built-in stats, personalisationClient-side performance and flicker; cost at scale
Feature flags + product experimentationGrowthBook (open source), PostHog, Statsig, LaunchDarkly, Optimizely Feature ExperimentationServer-side tests, gradual rollouts, apps and back-end logicNeeds developers; QA across services
Warehouse-native analysisGrowthBook, Statsig (warehouse-native option), Eppo and othersAnalyse experiments on your own data (BigQuery, Snowflake, Databricks, Redshift) with your metric definitionsRequires a data team and reliable event data
Platform-native experimentsGoogle Ads experiments, Meta A/B tests, email platform split tests, Shopify appsEasy for channel-specific questionsLimited to that channel; platform's own attribution

How to choose: start from your questions and constraints — who builds variants (marketers or developers), where changes live (web pages, checkout logic, apps), traffic volume, data maturity, performance budget, privacy and data-residency needs, and cost. Run a proof-of-concept A/A test on your own site before committing: it checks integration, SRM and flicker.

Variance reduction: CUPED in plain language

Tests are slow because metrics are noisy. Much of that noise comes from differences between users that existed before the test — some people simply buy more than others. CUPED (Controlled-experiment Using Pre-Experiment Data), introduced by researchers at Microsoft in 2013, uses each user's pre-experiment behaviour to remove that predictable part of the noise.

The idea in one line: adjusted metric = in-test metric − θ × (pre-test metric − average pre-test metric), where θ is chosen to minimise variance (θ = cov(pre, post) / var(pre)).

Because users are randomised, the pre-test metric has the same average in both groups, so the adjustment does not bias the difference between variants — it just makes it less noisy. The stronger the correlation between pre- and in-test behaviour, the bigger the variance reduction, and the smaller the sample you need. GrowthBook, Statsig, Optimizely and others offer CUPED or similar regression adjustment; check how each defines the pre-period.

Limits: CUPED helps little for new users with no history (common in ecommerce acquisition tests), and it does not fix peeking, SRM or bad metrics.

Hands-on: see CUPED work (simulation, standard library)

import random
from statistics import mean, variance, covariance   # Python 3.10+

random.seed(7)
users = []
for i in range(20_000):
    habit = random.expovariate(1 / 40)               # some users spend more, before and during
    pre = max(0.0, habit + random.gauss(0, 15))      # spend in 14 days before the test
    group = "B" if i % 2 else "A"
    post = max(0.0, habit + (1.5 if group == "B" else 0) + random.gauss(0, 15))  # true lift +1.5
    users.append((group, pre, post))

pres, posts = [u[1] for u in users], [u[2] for u in users]
theta, pre_mean = covariance(pres, posts) / variance(pres), mean(pres)

def diff_and_ci(adjust: bool):
    stats = {}
    for g in ("A", "B"):
        v = [post - theta * (pre - pre_mean) if adjust else post for grp, pre, post in users if grp == g]
        stats[g] = (mean(v), variance(v), len(v))
    d = stats["B"][0] - stats["A"][0]
    se = (stats["A"][1] / stats["A"][2] + stats["B"][1] / stats["B"][2]) ** 0.5
    return d, 1.96 * se

for label, adj in (("Raw", False), ("CUPED", True)):
    d, ci = diff_and_ci(adj)
    print(f"{label:5s} difference {d:5.2f} ± {ci:.2f}")

In one run (seed 7) the raw estimate was about 1.93 ± 1.14 and the CUPED estimate about 1.14 ± 0.52: the interval is less than half as wide, which is equivalent to needing far fewer users for the same precision. Your output will vary with the seed; the point is the narrower interval, not the exact numbers.

Other ways to speed up learning

  • Better metrics: a sensitive, high-frequency metric (e.g. add-to-cart) as a secondary signal alongside the primary.
  • Capping (winsorising) extreme revenue values, decided before launch.
  • Bolder changes: bigger expected effects need smaller samples.
  • Holdouts and long-term measurement for cumulative impact (see Measuring CRO impact).

Where AI fits in experimentation platforms

Vendors increasingly add AI to generate variant copy, suggest targeting and summarise results; some expose MCP servers or APIs so assistants can create and read experiments. Useful — as long as statistical rules, pre-registration and human review stay in place. An AI-written summary of a test must include the interval, SRM status and guardrails, not just "variant B won".

Common mistakes

  • Choosing a platform before defining questions, owners and data flows.
  • Skipping an A/A test when adopting a new tool.
  • Expecting CUPED to help much in tests dominated by new visitors.
  • Switching statistical settings mid-test.
  • Letting AI summaries replace the full test report.

Key takeaways

  • Google Optimize was sunset on 30 September 2023; the market now spans marketer-led web testing, feature flags, warehouse-native analysis and platform-native experiments.
  • Choose a stack from your questions, team, traffic, data maturity, performance and privacy needs — then validate it with an A/A test.
  • CUPED uses pre-experiment behaviour to remove predictable noise without biasing the treatment effect, shrinking the sample needed.
  • CUPED helps most for returning users with history and little for brand-new visitors; it does not fix peeking, SRM or bad metrics.
  • AI features in platforms are useful, but reports must still include intervals, SRM checks and guardrails.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why is an A/A test a good first experiment on a new testing platform?
  2. In which test would CUPED usually help least?
  3. Which statement about CUPED is correct?

Put it into practice

List your top three experimentation questions for next quarter, map each to a tool category, and plan an A/A test on your current setup with pass criteria for split, data match, flicker and speed.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.