AI-Powered Performance MarketingAttribution, incrementality and mix modeling · Lesson 11 of 15

Incrementality testing: geo experiments, conversion lift and holdouts

Article · 8 min · 8 min lecture

Video lecture

Incrementality testing: geo experiments, conversion lift and holdouts

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Incrementality testing

  • The one question that matters
  • Three designs
  • Geo test step by step
  • Power and decisions

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The only question that matters

Incrementality asks: how many more conversions (or how much more profit) happened because of this advertising, compared with what would have happened without it? Every other metric is a proxy.

Three test designs

DesignHow it worksStrengthsLimitations
User-level lift (conversion lift / brand lift)Platform randomly withholds ads from a control group of users and compares conversion ratesRandomized, run inside the platformPlatform runs and grades the test; needs enough conversions; limited to one platform
Geo experimentsRegions are assigned to treatment (ads on/increased) or control (ads off/unchanged); compare outcomes using your own sales dataUses your backend data; works across channels and offline salesNeeds enough regions and stable regional data; spillover between regions
Holdouts / on-off testsA share of your customer list or a region is excluded for an extended periodSimple, good for retargeting, CRM and emailOpportunity cost; long duration

Platforms offer lift tooling: Meta Conversion Lift, Google Ads Conversion Lift (Google has moved to a Bayesian methodology that allows smaller studies than before — check current eligibility thresholds in your account), and TikTok lift studies (availability varies). For geo tests, open-source tools include Meta's GeoLift (R) and approaches from Google's research (such as time-based regression and trimmed match). Meridian and Robyn can also be calibrated using geo test results.

Designing a geo test

  1. Question: "Does our YouTube Demand Gen spend increase total orders in the UK?"
  2. Unit: regions (for example UK postcode areas or DMA-like regions; in the Gulf, emirates or cities; in Pakistan, cities).
  3. Pre-period: 8–12+ weeks of historical data to confirm treatment and control move together (parallel trends).
  4. Assignment: pick matched markets using a tool (GeoLift's market selection) rather than gut feel.
  5. Treatment: turn spend on/off or scale it by a meaningful amount (small changes produce undetectable effects).
  6. Duration: long enough to capture the conversion lag; often 3–6 weeks plus a cool-down.
  7. Primary metric: backend orders or revenue by region — not platform conversions.
  8. Analysis: synthetic control or difference-in-differences; report the lift with an interval, and compute incremental CPA or iROAS.
Incremental conversions = actual treatment conversions − predicted conversions without ads
iROAS = incremental revenue ÷ test spend
Incremental CPA = test spend ÷ incremental conversions

Power: can your test even detect the effect?

A test that cannot detect a plausible effect wastes money. Before launching, simulate: "If the true lift is 5%, would this design detect it?" GeoLift includes power analysis. If the answer is no, increase the spend change, add regions, lengthen the test, or accept a directional result.

Worked example: a Dubai e-commerce retailer

A UAE electronics retailer wonders whether Meta prospecting adds sales beyond Google Search. Design (illustrative):

  • Units: emirates and major city zones; Dubai and Abu Dhabi split into matched zones by delivery postcode.
  • Treatment: Meta prospecting paused in control zones for four weeks; all other channels unchanged.
  • Metric: backend orders by delivery zone.
  • Result: an estimated lift with a confidence interval; incremental CPA was about 1.4× the platform-reported CPA (illustrative), giving an incrementality factor of about 0.7. Meta stayed, with a lower budget target, and a Conversion Lift study was scheduled to cross-check.

Hands-on: GeoLift skeleton (R)

# install.packages("remotes"); remotes::install_github("facebookincubator/GeoLift")
library(GeoLift)

data <- read.csv("orders_by_region_daily.csv")  # columns: date, location, Y (orders)
geo <- GeoDataRead(data = data, date_id = "date", location_id = "location",
                   Y_id = "Y", format = "yyyy-mm-dd")

# Power analysis and market selection
markets <- GeoLiftMarketSelection(data = geo, treatment_periods = c(28),
                                  N = c(3, 4, 5), effect_size = seq(0, 0.2, 0.05),
                                  lookback_window = 7, include_markets = c(),
                                  exclude_markets = c(), cpic = 25, budget = 20000,
                                  alpha = 0.1, fixed_effects = TRUE)
print(markets$BestMarkets)

# After the test: estimate lift
# result <- GeoLift(Y_id = "Y", data = geo_test, locations = c("zone_a","zone_b"),
#                   treatment_start_time = 91, treatment_end_time = 118)
# summary(result)

Check the GeoLift documentation for current function arguments; they change between versions. cpic is your estimated cost per incremental conversion.

Pitfalls

  • Using platform conversions as the outcome in a geo test.
  • Too few regions or too small a spend change.
  • Running other big changes (promotions, price changes) in only some regions.
  • Stopping early when the result looks good.

How to measure success

A test calendar (at least one causal test per major channel per half-year), documented results with intervals, and budget decisions that reference incrementality factors.

Key takeaways

  • Incrementality measures what happened because of ads versus without them — everything else is a proxy.
  • Use platform lift studies, geo experiments and holdouts, each with known strengths and limits.
  • Design geo tests with matched markets, pre-period checks, meaningful treatment, backend outcomes and power analysis.
  • Report incremental CPA or iROAS with intervals and convert results into incrementality factors.
  • Keep a test calendar covering every major channel.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. In a geo test, what should be the primary outcome metric?
  2. A power analysis shows your design can detect only a 20% lift, but you expect about 5%. What should you do?
  3. Test spend was 50,000 and incremental revenue 150,000. What is iROAS?

Put it into practice

Design a geo test for one channel: question, regions, pre-period, treatment, duration, primary metric and the minimum detectable effect you need.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.