---
title: "Incrementality testing: geo experiments, conversion lift…"
description: "The only question that matters Incrementality asks: how many more conversions (or how much more profit) happened because of this advertising, compared…"
url: https://optimizeall.com/learn/ai-performance-marketing/incrementality-testing
updated: 2026-10-05
---

AI-Powered Performance Marketing · Attribution, incrementality and mix modeling · lesson 11 of 15 · 8 min

# Incrementality testing: geo experiments, conversion lift and holdouts

## The only question that matters

Incrementality asks: **how many more conversions (or how much more profit) happened because of this advertising, compared with what would have happened without it?** Every other metric is a proxy.

## Three test designs

| Design | How it works | Strengths | Limitations |
|---|---|---|---|
| **User-level lift (conversion lift / brand lift)** | Platform randomly withholds ads from a control group of users and compares conversion rates | Randomized, run inside the platform | Platform runs and grades the test; needs enough conversions; limited to one platform |
| **Geo experiments** | Regions are assigned to treatment (ads on/increased) or control (ads off/unchanged); compare outcomes using your own sales data | Uses your backend data; works across channels and offline sales | Needs enough regions and stable regional data; spillover between regions |
| **Holdouts / on-off tests** | A share of your customer list or a region is excluded for an extended period | Simple, good for retargeting, CRM and email | Opportunity cost; long duration |

Platforms offer lift tooling: Meta Conversion Lift, Google Ads Conversion Lift (Google has moved to a Bayesian methodology that allows smaller studies than before — check current eligibility thresholds in your account), and TikTok lift studies (availability varies). For geo tests, open-source tools include Meta's **GeoLift** (R) and approaches from Google's research (such as time-based regression and trimmed match). Meridian and Robyn can also be *calibrated* using geo test results.

## Designing a geo test

1. **Question**: "Does our YouTube Demand Gen spend increase total orders in the UK?"
2. **Unit**: regions (for example UK postcode areas or DMA-like regions; in the Gulf, emirates or cities; in Pakistan, cities).
3. **Pre-period**: 8–12+ weeks of historical data to confirm treatment and control move together (parallel trends).
4. **Assignment**: pick matched markets using a tool (GeoLift's market selection) rather than gut feel.
5. **Treatment**: turn spend on/off or scale it by a meaningful amount (small changes produce undetectable effects).
6. **Duration**: long enough to capture the conversion lag; often 3–6 weeks plus a cool-down.
7. **Primary metric**: backend orders or revenue by region — not platform conversions.
8. **Analysis**: synthetic control or difference-in-differences; report the lift with an interval, and compute **incremental CPA** or **iROAS**.

```text
Incremental conversions = actual treatment conversions − predicted conversions without ads
iROAS = incremental revenue ÷ test spend
Incremental CPA = test spend ÷ incremental conversions
```

## Power: can your test even detect the effect?

A test that cannot detect a plausible effect wastes money. Before launching, simulate: "If the true lift is 5%, would this design detect it?" GeoLift includes power analysis. If the answer is no, increase the spend change, add regions, lengthen the test, or accept a directional result.

## Worked example: a Dubai e-commerce retailer

A UAE electronics retailer wonders whether Meta prospecting adds sales beyond Google Search. Design (illustrative):

- Units: emirates and major city zones; Dubai and Abu Dhabi split into matched zones by delivery postcode.
- Treatment: Meta prospecting paused in control zones for four weeks; all other channels unchanged.
- Metric: backend orders by delivery zone.
- Result: an estimated lift with a confidence interval; incremental CPA was about 1.4× the platform-reported CPA (illustrative), giving an incrementality factor of about 0.7. Meta stayed, with a lower budget target, and a Conversion Lift study was scheduled to cross-check.

## Hands-on: GeoLift skeleton (R)

```r
# install.packages("remotes"); remotes::install_github("facebookincubator/GeoLift")
library(GeoLift)

data <- read.csv("orders_by_region_daily.csv")  # columns: date, location, Y (orders)
geo <- GeoDataRead(data = data, date_id = "date", location_id = "location",
                   Y_id = "Y", format = "yyyy-mm-dd")

# Power analysis and market selection
markets <- GeoLiftMarketSelection(data = geo, treatment_periods = c(28),
                                  N = c(3, 4, 5), effect_size = seq(0, 0.2, 0.05),
                                  lookback_window = 7, include_markets = c(),
                                  exclude_markets = c(), cpic = 25, budget = 20000,
                                  alpha = 0.1, fixed_effects = TRUE)
print(markets$BestMarkets)

# After the test: estimate lift
# result <- GeoLift(Y_id = "Y", data = geo_test, locations = c("zone_a","zone_b"),
#                   treatment_start_time = 91, treatment_end_time = 118)
# summary(result)
```

Check the GeoLift documentation for current function arguments; they change between versions. `cpic` is your estimated cost per incremental conversion.

## Pitfalls

- Using platform conversions as the outcome in a geo test.
- Too few regions or too small a spend change.
- Running other big changes (promotions, price changes) in only some regions.
- Stopping early when the result looks good.

## How to measure success

A test calendar (at least one causal test per major channel per half-year), documented results with intervals, and budget decisions that reference incrementality factors.

## Video lecture: Incrementality testing: geo experiments, conversion lift and holdouts

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. Incrementality testing
2. Why it matters
3. User-level lift
4. Geo experiments and holdouts
5. Geo test, part one
6. Simple example: Lahore grocer holdout (illustrative)
7. Geo test, part two
8. Power analysis
9. Example: UAE electronics (illustrative)
10. Four mistakes
11. Mistakes + try this now
12. Watch me do it: YouTube geo test (illustrative)
13. Recap and next step

## Lecture transcript

### Incrementality testing

Every metric we've discussed so far is, in the end, a proxy. There's only one question that really matters: how much more did you sell because of this advertising, compared with what would have happened without it? That's incrementality. In this lecture you'll learn the three main test designs, how to design a geo experiment step by step, why power analysis saves you from wasted tests, and how to turn results into decisions.

### Why it matters

Why does this matter? Because without incrementality, you can't tell the difference between advertising that creates sales and advertising that just stands next to sales. Here's an analogy. A rooster crows every morning, and every morning the sun rises. If you only looked at the correlation, you'd credit the rooster. Plenty of ad spend is a rooster: it shows up right before people buy, but they'd have bought anyway. An incrementality test is simply putting the rooster in another room for a few days and seeing if the sun still rises.

### User-level lift

Design one: user-level lift. The platform randomly holds back ads from a control group and compares conversion rates. It's properly randomized and easy to run. The catch is that the platform runs and grades its own test, you need enough conversions, and it covers one platform at a time. Meta and Google both offer Conversion Lift, and Google has moved to a Bayesian method that allows smaller studies than before. Check current eligibility in your account.

### Geo experiments and holdouts

Design two: geo experiments. You split regions into treatment, where ads run or increase, and control, where they're off or unchanged. Then you compare your own sales by region. The big advantage: it uses your backend data, so it captures offline sales and works across channels. The challenges: you need enough regions with stable data, and people move between regions. Design three: holdouts, where part of your customer list or a region gets no ads for a longer period. Great for retargeting and email.

### Geo test, part one

Let's design a geo test. Start with a crisp question, like: does our YouTube spend increase total orders in the UK? Choose your units, maybe postcode areas in the UK, emirates or city zones in the Gulf, cities in Pakistan. Gather eight to twelve weeks or more of history to confirm treatment and control move together. Then use a tool, not gut feel, to select matched markets. Meta's open-source GeoLift does this.

### Simple example: Lahore grocer holdout (illustrative)

Let's walk a simple worked example. An online grocer in Lahore wonders if its retargeting ads help. They create a holdout: ten percent of recent site visitors, chosen at random, are excluded from retargeting for four weeks. At the end, the retargeted group bought at a rate of eight percent, the holdout at seven point five percent. So only half a percentage point is incremental. If retargeting cost fifty thousand rupees and the extra purchases earned about the same in margin, it's barely breaking even, even though the platform reported a lovely return. These numbers are illustrative, but holdouts like this regularly reveal it.

### Geo test, part two

Part two. Make the treatment meaningful. Tiny spend changes produce effects too small to see. Run long enough to capture the lag between seeing an ad and buying, often three to six weeks, plus a cool-down. Measure backend orders or revenue by region, never platform conversions. Analyze with a synthetic control or difference-in-differences, report the lift with an interval, and calculate incremental cost per acquisition and incremental return on ad spend.

### Power analysis

Now power, the step people skip. Before you spend a single dirham, ask: if the true lift were five percent, would this design detect it? If the answer is no, you're about to buy an inconclusive result. Fix it by making a bigger spend change, adding regions, or running longer. Or consciously accept a directional result. GeoLift includes power analysis, and the lesson has a skeleton in R to get you started.

### Example: UAE electronics (illustrative)

Here's an illustrative example. A UAE electronics retailer asked whether Meta prospecting adds sales beyond Google Search. They matched zones by delivery postcode across Dubai and Abu Dhabi, paused Meta prospecting in control zones for four weeks, and measured backend orders by zone. The incremental cost per acquisition came out about one point four times the platform-reported figure, giving an incrementality factor of roughly zero point seven. Meta stayed, but with a revised target and a Conversion Lift study scheduled to cross-check.

### Four mistakes

Avoid four classic mistakes. Using platform conversions as your geo test outcome. Too few regions or too small a change. Running promotions or price changes in only some regions, which contaminates the test. And stopping early because the result looks good. Build a test calendar instead, with at least one causal test per major channel every six months, and store every result with its interval.

### Mistakes + try this now

Common incrementality mistakes. Using platform-reported conversions as the outcome in a geo test. Choosing test regions by gut feel instead of matching them. Running a promotion in only some regions during the test. Making the spend change so small that no effect could ever be detected. And stopping as soon as the result looks nice. Try this now: pick one channel you're unsure about and write a one-paragraph test design: the question, treatment and control, the outcome from your backend, the duration, and your decision rule.

### Watch me do it: YouTube geo test (illustrative)

Watch me do it. Let's design an illustrative geo test for a UK online bakery on YouTube. Question: does YouTube Demand Gen increase total orders in England? Step one, units: I group delivery postcodes into forty areas with stable weekly orders. Step two, history: twelve weeks of orders by area. Step three, market selection: I run GeoLift's market selection, asking for treatment groups of five areas, a twenty-eight-day test, and our budget. It returns candidate sets with their power: the best set can detect roughly a five percent lift. Step four, treatment: YouTube runs only in the five treatment areas, at a meaningful budget, while everything else stays the same everywhere. Step five, contamination check: the promotions calendar shows a discount planned nationally in week three. That's fine because it applies to all areas equally. Step six, decision rule, written before launch: if incremental cost per order is below our target, scale YouTube nationally; if above, cut back; if inconclusive, repeat with a larger budget. Step seven, after the test, analyze with GeoLift and report the lift with its interval, not just a single number.

### Recap and next step

Recap. Incrementality is the question that matters. Use lift studies, geo tests and holdouts for their strengths. Design geo tests with matched markets, meaningful treatment, backend outcomes and power analysis. Report incremental CPA or iROAS with intervals and turn them into incrementality factors. Your next step: design a geo test for one channel, including the question, regions, treatment, duration, metric and the minimum effect you need to detect.

## Key takeaways

- Incrementality measures what happened because of ads versus without them — everything else is a proxy.
- Use platform lift studies, geo experiments and holdouts, each with known strengths and limits.
- Design geo tests with matched markets, pre-period checks, meaningful treatment, backend outcomes and power analysis.
- Report incremental CPA or iROAS with intervals and convert results into incrementality factors.
- Keep a test calendar covering every major channel.

## Try it

Design a geo test for one channel: question, regions, pre-period, treatment, duration, primary metric and the minimum detectable effect you need.

- [Previous: Attribution reality: what each number can and cannot tell you](https://optimizeall.com/learn/ai-performance-marketing/attribution-reality)
- [Next: Marketing mix modeling with Meridian and Robyn](https://optimizeall.com/learn/ai-performance-marketing/marketing-mix-modelling)
- [All lessons of AI-Powered Performance Marketing](https://optimizeall.com/learn/ai-performance-marketing)
