Conversion Rate Optimization (CRO)A/B testing statistics and experimentation platforms · Lesson 17 of 20
Experimentation platforms, CUPED and variance reduction
Video lecture
Experimentation platforms, CUPED and variance reduction
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Experimentation platforms and CUPED
When Google Optimize shut down in September twenty twenty-three, thousands of teams had to answer a question they'd never really thought about: what should our experimentation stack actually be? In this lecture you'll learn the four categories of tools that replaced it, how to choose between them, why an A/A test is the best first experiment on any new platform, and a technique called CUPED that can make tests dramatically more precise using data you already have. Then you'll watch me simulate CUPED and see the confidence interval shrink.
0:39 Why the stack matters
Why does the stack matter? Because the tool shapes what you can test, who can build tests, how fast pages load, and whether you can trust the numbers. A marketer-friendly visual editor is perfect for copy and layout, but can't safely test checkout logic. A developer-focused feature-flag system handles pricing and apps, but marketers can't use it alone. And if your results live only in a vendor dashboard, you can't easily verify them. The right stack fits your questions, your team and your data.
1:16 Four categories
Here are the four categories. One: marketer-led web experimentation, like VWO, Optimizely Web Experimentation, AB Tasty, Kameleoon and Convert, with visual editors, targeting and built-in statistics. Two: feature flags and product experimentation, like GrowthBook, PostHog, Statsig and LaunchDarkly, for server-side tests and gradual rollouts. Three: warehouse-native analysis, where experiments are analysed on your own data in BigQuery, Snowflake or similar, using your own metric definitions. And four: platform-native experiments inside Google Ads, Meta, email tools and shop apps, for channel-specific questions.
1:51 Choosing a platform
How do you choose? Start from questions and constraints, not features. Who will build variants, marketers or developers? Where do changes live: web pages, checkout logic, mobile apps? How much traffic do you have? How mature is your data: do you have a warehouse and reliable events? What's your performance budget for extra scripts? What are your privacy and data-residency needs, for example for customers in the EU or the Gulf? And what can you afford at your traffic level? Many teams combine two categories: a web tool for marketing tests, and feature flags for product changes.
2:33 Validate with an A/A test
Before you trust any new platform, run an A/A test. Both groups get the identical experience. You're not testing a change. You're testing the tool. Does the split come out even, with no sample ratio mismatch? Do the conversion numbers match your analytics and backend? Is there flicker? Does it slow the page? In an A/A test, a significant difference should appear only about as often as your significance level predicts. If it shows up far more, or the split is off, fix the integration before running real tests.
3:12 Variance reduction
Now the most useful statistical idea you may not know yet: variance reduction. Tests are slow because metrics are noisy. And a lot of that noise comes from differences between people that existed before the test began. Some customers simply buy more than others. CUPED, which stands for controlled-experiment using pre-experiment data, and was introduced by researchers at Microsoft in twenty thirteen, uses each user's behaviour before the test to remove that predictable part of the noise.
3:45 The runner analogy
Here's the intuition, with an analogy. Imagine judging whether a new training plan makes runners faster. Some runners were already fast. If you compare raw race times, the natural differences drown out the effect of the plan. But if you compare each runner's time with their own previous best, the plan's effect stands out. CUPED does that for your metrics. Adjusted metric equals in-test metric minus theta times the difference between the user's pre-test metric and the average. Because users were randomised, this doesn't bias the difference between variants. It just removes predictable noise.
4:26 When CUPED helps
When does CUPED help, and when doesn't it? It helps most when behaviour before the test strongly predicts behaviour during it: returning customers, subscribers, logged-in users, spend or engagement metrics. It helps little for brand-new visitors with no history, which is common in acquisition tests for online stores. And it doesn't fix peeking, sample ratio mismatch or a bad primary metric. Many platforms, including GrowthBook, Statsig and Optimizely, offer CUPED or similar regression adjustment. Check how each defines the pre-experiment window.
5:01 Simple example: UK coffee subscriptions (illustrative)
A simple example. A subscription coffee brand in the UK tests a new account page designed to increase add-on purchases. Almost all users are existing subscribers with months of history. Without variance reduction, the test would need about eight weeks. With CUPED, using each subscriber's add-on spend in the previous four weeks, the required duration drops substantially, because past spend strongly predicts future spend. These are illustrative figures, but the pattern is typical for returning-customer tests.
5:34 Realistic example: UAE marketplace (illustrative)
Now a realistic scenario with illustrative details. A fashion marketplace in the UAE adopts a warehouse-native experimentation setup. Events from their app and website land in their data warehouse. Feature flags assign users server-side. Metrics are defined once, in SQL, and reused across every test. Their first experiment is an A/A test, which reveals that one app version wasn't sending exposure events, causing a mismatch. After the fix, they enable CUPED for returning-customer tests, and keep it off for tests aimed at first-time visitors, where it wouldn't help. They also agree that any AI-written test summary must include the interval, SRM status and guardrails.
6:19 Watch me do it: CUPED simulation
Watch me simulate CUPED. I generate twenty thousand users. Each has a spending habit, which drives both their spend in the fourteen days before the test and their spend during it. Variant B gets a small true lift. I compute theta as the covariance of pre and post spend divided by the variance of pre spend. Then I compare the difference between variants two ways: raw, and CUPED-adjusted. In one run, the raw estimate is about one point nine, plus or minus one point one. The CUPED estimate is about one point one, plus or minus zero point five. The interval is less than half as wide. That's like getting a much bigger sample for free.
7:09 AI in experimentation platforms
AI is arriving in experimentation platforms too. Vendors add AI to generate variant copy, suggest targeting and summarise results, and some expose APIs or MCP servers so assistants can create and read experiments. That's useful. But keep the rules in place: pre-registered hypotheses and metrics, the statistical method as designed, SRM checks, guardrails and human review. An AI summary that says variant B won, without the interval, the sample ratio check and the guardrails, is not a report. It's a headline.
7:44 Common mistakes
Common mistakes. Choosing a platform before defining questions, owners and data flows. Skipping the A/A test. Relying only on a vendor dashboard you can't verify. Expecting CUPED to rescue tests dominated by new visitors. Changing statistical settings mid-test. And letting AI summaries replace the full report.
8:04 Recap
Recap. After Google Optimize, the stack falls into four categories: marketer-led web testing, feature flags and product experimentation, warehouse-native analysis, and platform-native experiments. Choose from your questions, team, traffic, data and privacy needs, and validate any new tool with an A/A test. Use CUPED to reduce noise when pre-test behaviour predicts in-test behaviour, and remember its limits. Keep AI in a supporting role. Try this now: write down your top three experimentation questions for the next quarter, map each to a tool category, and plan an A/A test for your current setup.
The experimentation stack after Google Optimize
Google Optimize and Optimize 360 were sunset on 30 September 2023. Google did not launch a direct replacement; instead GA4 integrates with third-party testing tools. The market now looks roughly like this (features change often — check vendors' current documentation and pricing):
| Category | Examples | Typical strengths | Consider |
|---|---|---|---|
| Marketer-led web experimentation | VWO, Optimizely (Web Experimentation), AB Tasty, Kameleoon, Convert | Visual editors, targeting, built-in stats, personalisation | Client-side performance and flicker; cost at scale |
| Feature flags + product experimentation | GrowthBook (open source), PostHog, Statsig, LaunchDarkly, Optimizely Feature Experimentation | Server-side tests, gradual rollouts, apps and back-end logic | Needs developers; QA across services |
| Warehouse-native analysis | GrowthBook, Statsig (warehouse-native option), Eppo and others | Analyse experiments on your own data (BigQuery, Snowflake, Databricks, Redshift) with your metric definitions | Requires a data team and reliable event data |
| Platform-native experiments | Google Ads experiments, Meta A/B tests, email platform split tests, Shopify apps | Easy for channel-specific questions | Limited to that channel; platform's own attribution |
How to choose: start from your questions and constraints — who builds variants (marketers or developers), where changes live (web pages, checkout logic, apps), traffic volume, data maturity, performance budget, privacy and data-residency needs, and cost. Run a proof-of-concept A/A test on your own site before committing: it checks integration, SRM and flicker.
Variance reduction: CUPED in plain language
Tests are slow because metrics are noisy. Much of that noise comes from differences between users that existed before the test — some people simply buy more than others. CUPED (Controlled-experiment Using Pre-Experiment Data), introduced by researchers at Microsoft in 2013, uses each user's pre-experiment behaviour to remove that predictable part of the noise.
The idea in one line: adjusted metric = in-test metric − θ × (pre-test metric − average pre-test metric), where θ is chosen to minimise variance (θ = cov(pre, post) / var(pre)).
Because users are randomised, the pre-test metric has the same average in both groups, so the adjustment does not bias the difference between variants — it just makes it less noisy. The stronger the correlation between pre- and in-test behaviour, the bigger the variance reduction, and the smaller the sample you need. GrowthBook, Statsig, Optimizely and others offer CUPED or similar regression adjustment; check how each defines the pre-period.
Limits: CUPED helps little for new users with no history (common in ecommerce acquisition tests), and it does not fix peeking, SRM or bad metrics.
Hands-on: see CUPED work (simulation, standard library)
import random
from statistics import mean, variance, covariance # Python 3.10+
random.seed(7)
users = []
for i in range(20_000):
habit = random.expovariate(1 / 40) # some users spend more, before and during
pre = max(0.0, habit + random.gauss(0, 15)) # spend in 14 days before the test
group = "B" if i % 2 else "A"
post = max(0.0, habit + (1.5 if group == "B" else 0) + random.gauss(0, 15)) # true lift +1.5
users.append((group, pre, post))
pres, posts = [u[1] for u in users], [u[2] for u in users]
theta, pre_mean = covariance(pres, posts) / variance(pres), mean(pres)
def diff_and_ci(adjust: bool):
stats = {}
for g in ("A", "B"):
v = [post - theta * (pre - pre_mean) if adjust else post for grp, pre, post in users if grp == g]
stats[g] = (mean(v), variance(v), len(v))
d = stats["B"][0] - stats["A"][0]
se = (stats["A"][1] / stats["A"][2] + stats["B"][1] / stats["B"][2]) ** 0.5
return d, 1.96 * se
for label, adj in (("Raw", False), ("CUPED", True)):
d, ci = diff_and_ci(adj)
print(f"{label:5s} difference {d:5.2f} ± {ci:.2f}")In one run (seed 7) the raw estimate was about 1.93 ± 1.14 and the CUPED estimate about 1.14 ± 0.52: the interval is less than half as wide, which is equivalent to needing far fewer users for the same precision. Your output will vary with the seed; the point is the narrower interval, not the exact numbers.
Other ways to speed up learning
- Better metrics: a sensitive, high-frequency metric (e.g. add-to-cart) as a secondary signal alongside the primary.
- Capping (winsorising) extreme revenue values, decided before launch.
- Bolder changes: bigger expected effects need smaller samples.
- Holdouts and long-term measurement for cumulative impact (see Measuring CRO impact).
Where AI fits in experimentation platforms
Vendors increasingly add AI to generate variant copy, suggest targeting and summarise results; some expose MCP servers or APIs so assistants can create and read experiments. Useful — as long as statistical rules, pre-registration and human review stay in place. An AI-written summary of a test must include the interval, SRM status and guardrails, not just "variant B won".
Common mistakes
- Choosing a platform before defining questions, owners and data flows.
- Skipping an A/A test when adopting a new tool.
- Expecting CUPED to help much in tests dominated by new visitors.
- Switching statistical settings mid-test.
- Letting AI summaries replace the full test report.
Key takeaways
- Google Optimize was sunset on 30 September 2023; the market now spans marketer-led web testing, feature flags, warehouse-native analysis and platform-native experiments.
- Choose a stack from your questions, team, traffic, data maturity, performance and privacy needs — then validate it with an A/A test.
- CUPED uses pre-experiment behaviour to remove predictable noise without biasing the treatment effect, shrinking the sample needed.
- CUPED helps most for returning users with history and little for brand-new visitors; it does not fix peeking, SRM or bad metrics.
- AI features in platforms are useful, but reports must still include intervals, SRM checks and guardrails.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
List your top three experimentation questions for next quarter, map each to a tool category, and plan an A/A test on your current setup with pass criteria for split, data match, flicker and speed.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.