Conversion Rate Optimization (CRO)A/B testing statistics and experimentation platforms · Lesson 14 of 20

A/B test fundamentals and experiment design

Article · 11 min · 9 min lecture

Video lecture

A/B test fundamentals and experiment design

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

A/B test fundamentals

  • Sound tests before sound maths
  • Types of tests and test plans
  • Randomisation units
  • SRM, flicker and QA

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What an A/B test does

An A/B test randomly splits visitors between a control (the current experience, A) and one or more variants (B, C…). Because assignment is random, the groups are statistically similar in everything except the change. Differences in outcomes can therefore be attributed to the change — within the bounds of statistical uncertainty.

Randomisation is what separates a test from a before/after comparison. Seasonality, campaigns, news and competitor moves affect both groups equally.

Types of tests

TypeWhat it doesWhen to use
A/BControl vs one variantMost tests
A/B/nControl vs several variantsWhen you have multiple strong ideas and enough traffic
Multivariate (MVT)Tests combinations of several elementsVery high traffic; interaction effects matter
Split URLVariants on different URLsLarge redesigns or different templates
Server-side / feature flagsChanges delivered from the server or app codePricing logic, algorithms, apps, performance-sensitive changes

Every additional variant splits traffic further and requires more total sample, so keep variants few unless traffic is large.

Designing a test: the plan

Write a test plan before building:

Test name:        PDP delivery estimate (mobile)
Hypothesis:       [from backlog]
Audience:         Mobile visitors to product pages, all countries served
Allocation:       50/50
Primary metric:   Revenue per visitor (or purchase conversion)
Secondary:        Add-to-cart rate, checkout starts
Guardrails:       Returns rate, page load time, support contacts about delivery
Minimum detectable effect: (see next lesson)
Sample size / duration:    (calculated), at least 2 full business cycles
Stop rules:       Stop early only for guardrail harm or broken experience
QA:               Devices, browsers, languages, tracking of all metrics

Unit of randomisation

Most web tests randomise by user (via a cookie or user ID) so each person sees a consistent experience. Randomising by page view would show someone different versions on each visit — confusing and statistically messy. For logged-in products, user-ID-based assignment works across devices.

Sample ratio mismatch (SRM)

If you allocate 50/50 but observe, for example, 52,000 visitors in A and 47,000 in B, something may be wrong: a redirect losing users, a bot filter affecting one variant, or a bug preventing the variant from loading on certain browsers. A sample ratio mismatch check (a simple chi-square test many tools run automatically) should be passed before trusting results. An SRM invalidates the test until explained.

The flicker effect

Client-side testing tools change the page in the browser after it starts loading. If the original briefly appears before the variant (flicker), results are biased and the experience suffers. Mitigations: load the testing script correctly (often synchronously in the head with anti-flicker handling), keep variant code light, or use server-side testing for major changes.

QA before launch

  • Variant renders correctly on key devices, browsers and languages (including right-to-left).
  • Tracking fires for all metrics in all variants.
  • Allocation works (check initial traffic split).
  • No conflicts with other running tests on the same pages.
  • Exclusions (internal traffic, bots) configured.

Worked example: a test that was invalid

An e-commerce brand tests a new product page and sees the variant "winning" strongly after two weeks. The SRM check shows the variant had noticeably fewer users than expected. Investigation reveals the variant broke on an older mobile browser, so those users — who converted poorly anyway — were excluded from the variant's data but not the control's. The "win" was an artefact. After fixing the bug and restarting, the effect was much smaller.

Hands-on: check for sample ratio mismatch yourself

Most platforms run an SRM check automatically, but you should be able to verify it from raw counts (for example from your data warehouse). This uses only the Python standard library:

from statistics import NormalDist

def srm_p_value(users_a: int, users_b: int, expected_share_a: float = 0.5) -> float:
    """Chi-square goodness-of-fit test (1 degree of freedom) for a two-way split."""
    total = users_a + users_b
    exp_a, exp_b = total * expected_share_a, total * (1 - expected_share_a)
    chi2 = (users_a - exp_a) ** 2 / exp_a + (users_b - exp_b) ** 2 / exp_b
    return 2 * (1 - NormalDist().cdf(chi2 ** 0.5))   # with 1 df, chi2 = z^2

for a, b in [(50_210, 49_790), (50_600, 49_400)]:
    p = srm_p_value(a, b)
    verdict = "SRM - investigate before reading results" if p < 0.001 else "OK"
    print(f"A={a:,} B={b:,}  p={p:.5f}  {verdict}")
# A=50,210 B=49,790  p=0.18413  OK
# A=50,600 B=49,400  p=0.00015  SRM - investigate before reading results

A strict threshold (commonly p < 0.001) is used because SRM checks are run on every test and a false alarm is costly; any SRM invalidates the results until explained. A 1.2% imbalance looks small but, on 100,000 users, is very unlikely to be chance.

Client-side, server-side and edge testing in 2026

ApproachHow it worksGood forWatch out for
Client-side (visual editor)JavaScript changes the page in the browserCopy, layout, content tests by marketersFlicker, page weight, consent-dependent loading, breaks after site releases
Server-side / feature flagsYour code decides which version to renderPricing, search, checkout logic, apps, performance-sensitive pagesNeeds developers; QA across services
Edge (CDN)Variant chosen at the CDN before HTML reaches the browserFast full-page variants without flickerSetup complexity; caching rules

Platforms such as VWO, Optimizely, AB Tasty, Kameleoon and Convert offer visual editors and server-side SDKs; GrowthBook, PostHog and Statsig are feature-flag and product-experimentation oriented. The next lessons cover statistics engines and variance reduction.

Where consent is required for non-essential cookies or storage, testing tools that set identifiers may need consent before they run, which changes who is in your experiment. Document how your setup handles non-consenting users (excluded, or served the control without tracking) and remember that results then describe consenting users.

Common mistakes

  • Starting tests without a written plan and primary metric.
  • Running overlapping tests on the same element without coordination.
  • Ignoring SRM warnings.
  • Launching variants without cross-device QA.
  • Using too many variants for the available traffic.

Choosing a testing approach

Client-side tools are quick to start and let marketers build simple variants, but they can add page weight and flicker. Server-side testing and feature flags require developer involvement but handle complex changes (pricing logic, search algorithms, app experiences) without flicker, and they suit high-traffic or performance-sensitive sites. Many mature programmes use both: client-side for content and layout tests, server-side for product and pricing changes. Whatever you choose, make sure analytics can see which variant each user received so results can be verified independently of the testing tool.

Key takeaways

  • Randomisation makes groups comparable, so differences can be attributed to the change.
  • Write a test plan with hypothesis, audience, metrics, sample size and stop rules.
  • Randomise by user and check for sample ratio mismatch before trusting results.
  • QA variants across devices, browsers and languages before launch.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A 50/50 test shows 60,000 users in control and 51,000 in the variant. What should you do?
  2. Why do most web tests randomise by user rather than by page view?
  3. What is the main risk of adding more variants to a test?
  4. Your 50/50 test shows 50,600 users in A and 49,400 in B, with an SRM p-value around 0.00015. What should you do?

Put it into practice

Write a complete test plan for one hypothesis from your backlog using the template in this lesson.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.