---
title: "A/B test fundamentals and experiment design"
description: "What an A/B test does An A/B test randomly splits visitors between a control (the current experience, A) and one or more variants (B, C…). Because…"
url: https://optimizeall.com/learn/conversion-rate-optimization/ab-test-fundamentals
updated: 2026-10-05
---

Conversion Rate Optimization (CRO) · A/B testing statistics and experimentation platforms · lesson 14 of 20 · 11 min

# A/B test fundamentals and experiment design

## What an A/B test does

An **A/B test** randomly splits visitors between a **control** (the current experience, A) and one or more **variants** (B, C…). Because assignment is random, the groups are statistically similar in everything except the change. Differences in outcomes can therefore be attributed to the change — within the bounds of statistical uncertainty.

Randomisation is what separates a test from a before/after comparison. Seasonality, campaigns, news and competitor moves affect both groups equally.

## Types of tests

| Type | What it does | When to use |
|---|---|---|
| A/B | Control vs one variant | Most tests |
| A/B/n | Control vs several variants | When you have multiple strong ideas and enough traffic |
| Multivariate (MVT) | Tests combinations of several elements | Very high traffic; interaction effects matter |
| Split URL | Variants on different URLs | Large redesigns or different templates |
| Server-side / feature flags | Changes delivered from the server or app code | Pricing logic, algorithms, apps, performance-sensitive changes |

Every additional variant splits traffic further and requires more total sample, so keep variants few unless traffic is large.

## Designing a test: the plan

Write a test plan before building:

```
Test name:        PDP delivery estimate (mobile)
Hypothesis:       [from backlog]
Audience:         Mobile visitors to product pages, all countries served
Allocation:       50/50
Primary metric:   Revenue per visitor (or purchase conversion)
Secondary:        Add-to-cart rate, checkout starts
Guardrails:       Returns rate, page load time, support contacts about delivery
Minimum detectable effect: (see next lesson)
Sample size / duration:    (calculated), at least 2 full business cycles
Stop rules:       Stop early only for guardrail harm or broken experience
QA:               Devices, browsers, languages, tracking of all metrics
```

## Unit of randomisation

Most web tests randomise by **user** (via a cookie or user ID) so each person sees a consistent experience. Randomising by page view would show someone different versions on each visit — confusing and statistically messy. For logged-in products, user-ID-based assignment works across devices.

## Sample ratio mismatch (SRM)

If you allocate 50/50 but observe, for example, 52,000 visitors in A and 47,000 in B, something may be wrong: a redirect losing users, a bot filter affecting one variant, or a bug preventing the variant from loading on certain browsers. A **sample ratio mismatch check** (a simple chi-square test many tools run automatically) should be passed before trusting results. An SRM invalidates the test until explained.

## The flicker effect

Client-side testing tools change the page in the browser after it starts loading. If the original briefly appears before the variant (**flicker**), results are biased and the experience suffers. Mitigations: load the testing script correctly (often synchronously in the head with anti-flicker handling), keep variant code light, or use server-side testing for major changes.

## QA before launch

- Variant renders correctly on key devices, browsers and languages (including right-to-left).
- Tracking fires for all metrics in all variants.
- Allocation works (check initial traffic split).
- No conflicts with other running tests on the same pages.
- Exclusions (internal traffic, bots) configured.

## Worked example: a test that was invalid

An e-commerce brand tests a new product page and sees the variant "winning" strongly after two weeks. The SRM check shows the variant had noticeably fewer users than expected. Investigation reveals the variant broke on an older mobile browser, so those users — who converted poorly anyway — were excluded from the variant's data but not the control's. The "win" was an artefact. After fixing the bug and restarting, the effect was much smaller.

## Hands-on: check for sample ratio mismatch yourself

Most platforms run an SRM check automatically, but you should be able to verify it from raw counts (for example from your data warehouse). This uses only the Python standard library:

```python
from statistics import NormalDist

def srm_p_value(users_a: int, users_b: int, expected_share_a: float = 0.5) -> float:
    """Chi-square goodness-of-fit test (1 degree of freedom) for a two-way split."""
    total = users_a + users_b
    exp_a, exp_b = total * expected_share_a, total * (1 - expected_share_a)
    chi2 = (users_a - exp_a) ** 2 / exp_a + (users_b - exp_b) ** 2 / exp_b
    return 2 * (1 - NormalDist().cdf(chi2 ** 0.5))   # with 1 df, chi2 = z^2

for a, b in [(50_210, 49_790), (50_600, 49_400)]:
    p = srm_p_value(a, b)
    verdict = "SRM - investigate before reading results" if p < 0.001 else "OK"
    print(f"A={a:,} B={b:,}  p={p:.5f}  {verdict}")
# A=50,210 B=49,790  p=0.18413  OK
# A=50,600 B=49,400  p=0.00015  SRM - investigate before reading results
```

A strict threshold (commonly p < 0.001) is used because SRM checks are run on every test and a false alarm is costly; any SRM invalidates the results until explained. A 1.2% imbalance looks small but, on 100,000 users, is very unlikely to be chance.

## Client-side, server-side and edge testing in 2026

| Approach | How it works | Good for | Watch out for |
|---|---|---|---|
| Client-side (visual editor) | JavaScript changes the page in the browser | Copy, layout, content tests by marketers | Flicker, page weight, consent-dependent loading, breaks after site releases |
| Server-side / feature flags | Your code decides which version to render | Pricing, search, checkout logic, apps, performance-sensitive pages | Needs developers; QA across services |
| Edge (CDN) | Variant chosen at the CDN before HTML reaches the browser | Fast full-page variants without flicker | Setup complexity; caching rules |

Platforms such as VWO, Optimizely, AB Tasty, Kameleoon and Convert offer visual editors and server-side SDKs; GrowthBook, PostHog and Statsig are feature-flag and product-experimentation oriented. The next lessons cover statistics engines and variance reduction.

## Consent and testing

Where consent is required for non-essential cookies or storage, testing tools that set identifiers may need consent before they run, which changes who is in your experiment. Document how your setup handles non-consenting users (excluded, or served the control without tracking) and remember that results then describe consenting users.

## Common mistakes

- Starting tests without a written plan and primary metric.
- Running overlapping tests on the same element without coordination.
- Ignoring SRM warnings.
- Launching variants without cross-device QA.
- Using too many variants for the available traffic.

## Choosing a testing approach

Client-side tools are quick to start and let marketers build simple variants, but they can add page weight and flicker. Server-side testing and feature flags require developer involvement but handle complex changes (pricing logic, search algorithms, app experiences) without flicker, and they suit high-traffic or performance-sensitive sites. Many mature programmes use both: client-side for content and layout tests, server-side for product and pricing changes. Whatever you choose, make sure analytics can see which variant each user received so results can be verified independently of the testing tool.

## Video lecture: A/B test fundamentals and experiment design

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. A/B test fundamentals
2. Why design matters
3. Types of tests
4. The test plan
5. Metrics in the plan
6. Unit of randomisation
7. Sample ratio mismatch (SRM)
8. Flicker and QA
9. Simple example: an invalid 'win'
10. Realistic example: Dubai travel checkout (illustrative)
11. Watch me do it: SRM check in Python
12. Choosing an approach
13. Common mistakes
14. Recap

## Lecture transcript

### A/B test fundamentals

A test says your new product page lifts conversion by twelve percent. Before you celebrate, can you prove the test itself was sound? Most bad decisions in experimentation don't come from wrong maths. They come from broken tests: uneven splits, flickering variants, missing tracking, overlapping experiments. In this lecture you'll learn what an A/B test really does, the types of tests, how to write a test plan, how to choose the unit of randomisation, and how to catch sample ratio mismatch and flicker. Then you'll watch me check a split with a few lines of Python.

### Why design matters

Why is design so important? Because an A/B test is a measuring instrument. If the instrument is broken, the measurement is meaningless, however impressive it looks. Randomisation is what makes the instrument work. It creates two groups that are the same in every way except the change, so seasonality, campaigns, news and competitor moves affect both equally. Break the randomisation, even slightly, and you're back to comparing apples with oranges.

### Types of tests

Types of tests. A/B: control versus one variant, used for most tests. A/B/n: several variants, which needs more traffic because each variant gets a smaller share. Multivariate tests combine several elements and need very high traffic. Split URL tests send people to different pages, useful for big redesigns. And server-side or feature-flag tests change logic in your code, which suits pricing, search, checkout logic and apps. Every extra variant splits traffic further, so keep variants few unless traffic is large.

### The test plan

Write the plan before you build. Test name, hypothesis and audience. Allocation, usually fifty-fifty. The primary metric, secondary metrics and guardrails. The minimum detectable effect, the sample size and duration, covering at least full weeks. Stop rules: stop early only for guardrail harm or a broken experience. And a QA checklist. The plan is your contract with future you. It stops you from changing the question after you've seen the answer.

### Metrics in the plan

Let's talk about metrics inside the plan, because the plan is where most tests are won or lost. Choose one primary metric that decides the outcome, usually revenue per visitor or purchase conversion for a store, qualified leads for a service. Add a few secondary metrics that explain how the change worked, like add-to-cart rate or checkout starts. And add guardrails that must not get worse: return rate, page speed, support contacts, unsubscribes. Then write the decision rule in one sentence, before launch. For example: ship B if revenue per visitor improves at our planned confidence and no guardrail worsens meaningfully. That single sentence prevents most post-test arguments.

### Unit of randomisation

Unit of randomisation. Most web tests randomise by user, using a cookie or a logged-in user ID, so each person sees a consistent experience every visit. Randomising by page view would show someone different versions on each click, which is confusing and statistically messy. For logged-in products, user-ID assignment works across devices. And think about consent: where consent is required before storing identifiers, the tool may only run for consenting users, so your results describe that group.

### Sample ratio mismatch (SRM)

Now sample ratio mismatch. If you allocate fifty-fifty but see fifty thousand six hundred users in A and forty-nine thousand four hundred in B, something may be wrong. A redirect losing users, a bot filter hitting one variant, or a bug stopping the variant loading on some browsers. A simple chi-square test tells you how likely that imbalance is by chance. If it's very unlikely, commonly a p-value below one in a thousand, the test is invalid until you explain and fix it. Don't read the results first. Check the split first.

### Flicker and QA

Two more threats. Flicker: client-side tools change the page after it starts loading, so the original can flash before the variant appears, which biases results and annoys users. Mitigate by loading the script correctly with anti-flicker handling, keeping variant code light, or testing server-side or at the edge for big changes. And QA: check every variant on key devices, browsers and languages, including right-to-left layouts; confirm tracking fires for all metrics; check the initial split; and make sure no other tests conflict on the same pages.

### Simple example: an invalid 'win'

A simple example of an invalid test. An online store tests a new product page, and after two weeks the variant is winning strongly. The SRM check fails: the variant has noticeably fewer users than expected. Investigation shows the variant breaks on an older mobile browser, so those users, who convert poorly anyway, dropped out of the variant's data but not the control's. The win was an artefact. After fixing the bug and restarting, the effect is much smaller.

### Realistic example: Dubai travel checkout (illustrative)

Now a realistic scenario with illustrative details. A travel company in Dubai runs a checkout test using a client-side tool. Results look flat, and the SRM check passes. But QA had missed something: the variant's layout broke in Arabic, where the right-to-left version pushed the payment button below a sticky banner. Arabic-language users in the variant struggled, dragging the overall result down. The team fixes the layout, adds right-to-left checks to their QA template, and reruns. The lesson: SRM catches missing users, but only QA catches broken experiences for users who are still there.

### Watch me do it: SRM check in Python

Watch me check a split. I have user counts per variant from our data warehouse. I write a tiny function using Python's standard library: it computes the chi-square statistic for the observed split against the expected fifty-fifty and converts it to a p-value. I run it on two tests. Fifty thousand two hundred and ten versus forty-nine thousand seven hundred and ninety gives a p-value of about zero point one eight. Fine. Fifty thousand six hundred versus forty-nine thousand four hundred gives about zero point zero zero zero one five. That's a mismatch. A one point two percent imbalance looks tiny, but on a hundred thousand users it's very unlikely to be chance. So before anyone opens the results dashboard, I start investigating.

### Choosing an approach

Choosing a testing approach. Client-side tools with visual editors let marketers launch copy and layout tests quickly, but they add page weight and flicker, and they can break when the site changes. Server-side testing and feature flags need developers, but handle pricing, search and checkout logic without flicker. Edge testing, at the CDN, serves full-page variants before HTML reaches the browser. Many mature programmes use both client-side and server-side. Whatever you choose, make sure your analytics can see which variant each user received, so you can verify results independently of the tool.

### Common mistakes

Common mistakes. No written plan. Overlapping tests on the same element. Ignoring SRM warnings. Skipping cross-device and right-to-left QA. Too many variants for the traffic. Letting consent settings silently change who's in the test without documenting it. And trusting the tool's dashboard without being able to verify assignments in your own data.

### Recap

Recap. An A/B test is a measuring instrument that works only if randomisation works. Write the plan first, randomise by user, check the split for sample ratio mismatch before reading results, prevent flicker, and QA every variant across devices and languages. Choose client-side, server-side or edge approaches to suit the change. Try this now: take your last test, pull the user counts per variant, run the SRM function from the lesson text, and add right-to-left and in-app browser checks to your QA template.

## Key takeaways

- Randomisation makes groups comparable, so differences can be attributed to the change.
- Write a test plan with hypothesis, audience, metrics, sample size and stop rules.
- Randomise by user and check for sample ratio mismatch before trusting results.
- QA variants across devices, browsers and languages before launch.

## Try it

Write a complete test plan for one hypothesis from your backlog using the template in this lesson.

- [Previous: Optimising forms and checkout](https://optimizeall.com/learn/conversion-rate-optimization/forms-and-checkout)
- [Next: Sample size, significance and power](https://optimizeall.com/learn/conversion-rate-optimization/sample-size-and-significance)
- [All lessons of Conversion Rate Optimization (CRO)](https://optimizeall.com/learn/conversion-rate-optimization)
