AI-Powered Performance MarketingCreative as targeting: production and testing · Lesson 8 of 15
Creative testing at scale: design, read-outs and fatigue
Video lecture
Creative testing at scale: design, read-outs and fatigue
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Creative testing at scale
Most advertisers think they're testing creative. What they're actually doing is uploading ten ads and letting the algorithm pick a favorite. That's not a test, and it teaches you almost nothing you can reuse. In this lecture you'll learn the three modes of creative testing, how to design a controlled test that gives a clear answer, how to read results honestly, and how to stay ahead of creative fatigue.
0:30 Why it matters
Why does this matter? Because testing is how you turn spend into knowledge. Without it, every winning ad is a lucky accident you can't repeat. Here's an analogy. A chef who changes the salt, the spice and the cooking time all at once, and then the dish tastes better, has no idea why. Next week they can't do it again. A chef who changes one thing at a time builds a recipe book. Creative testing is how you build your recipe book: principles about hooks, formats, offers and languages that you can reuse on every future campaign.
1:12 Allocation isn't testing
Why isn't the algorithm's choice a test? Because it allocates spend based on early predictions, so the first ad to get lucky gets more budget. And different ads end up shown to different people. So when ad A wins, you don't know if it's the ad or the audience it happened to get. The right approach is to use both: let the algorithm scale winners, and run structured tests to learn what makes a winner.
1:45 Three modes
There are three modes. Mode one, exploration in-campaign. You add new concepts to your automated campaign every week and judge them after a minimum spend. Mode two, the controlled A/B test, using the platform's experiment tools that split audiences cleanly, for big questions like a new offer or positioning. Mode three, the lift test, which asks whether a creative strategy causes more sales at all. You'll build those in the measurement module.
2:16 Six design decisions
Designing a clean A/B test takes six decisions. One hypothesis, for example, showing the price upfront increases purchases from new customers. One variable, price shown or not, with everything else identical. A primary metric chosen before launch. Enough conversions per cell. As a rough illustrative rule, fewer than fifty to a hundred conversions per cell rarely gives a clear answer for moderate effects. At least a full week. And a decision rule written down in advance.
2:49 Simple example: Riyadh abaya store (illustrative)
Here's a simple worked example. A Riyadh abaya store wants to know whether showing the price in the first frame helps. Hypothesis: showing the price upfront increases purchases from new customers. Variable: price in the first frame or not, same video otherwise. Metric: cost per purchase. They use Meta's A/B test tool, split evenly, run for two weeks, and set a rule: ship the price version if it's at least ten percent cheaper per purchase with the tool's confidence above eighty percent. Result: price upfront wins. New principle for the recipe book: in this market, price transparency helps.
3:32 Read results honestly
Reading results is where people fool themselves. Look at your primary metric first, not whichever metric happens to favor the ad you like. Check audience composition: did one version reach more existing customers, who convert anyway? Watch for novelty effects, where a new ad spikes and then fades. And write the learning as a principle, like problem-first hooks beat product-first hooks for new customers in Saudi Arabia, not just ad fourteen won.
4:03 Creative fatigue
Now fatigue. When frequency climbs, click-through falls, and cost per result creeps up on a stable audience, the concept is wearing out. Automated campaigns rotate ads, which helps, but every concept eventually saturates its audience, and the more you spend, the faster that happens. The fix is a pipeline: always have new concepts in testing, so a replacement is ready before your current winner fades.
4:31 Example rhythm: Jeddah delivery (illustrative)
Here's an illustrative weekly rhythm from a food delivery startup in Jeddah. Monday, four new concepts into each automated campaign. Thursday, keep concepts with enough spend and a cost per first order within about thirty percent of target, pause the rest. Once a month, one controlled test on a strategic question. They tested Arabic-first voiceover against English with Arabic captions. Arabic-first won decisively for first orders, and it became a standing principle. Once a quarter, a geo lift test on whether the channel is incremental.
5:08 Analyze by concept
To analyze at the concept level, you need discipline in naming. Use a convention like concept, angle, format, version, separated by double underscores. Then a short Python script, like the one in the lesson, can export your ad data, group by concept, and calculate cost per acquisition, hook rate and new customer share, filtered by a minimum spend so you don't crown a winner on tiny numbers.
5:37 Mistakes + try this now
The common testing mistakes. Declaring a winner after thirty conversions. Changing the hook and the offer and the format in the same test. Letting the platform's spend allocation stand in for a test. Peeking every day and stopping early when you like the result. And, most wasteful of all, never writing down what you learned. Try this now: open a document called creative principles. Write down three things you believe about your customers' response to creative. Then mark each one as tested or assumed. Your next tests should target the assumptions.
6:17 Watch me do it: reading a concept report (illustrative)
Watch me do it. Here's how I'd read an illustrative concept report for a Doha perfume brand after two weeks. I export ad-level data and run the concept script. Six concepts. The script filters out two with spend below our minimum; they're not losers, just unproven, so they get another week. Of the remaining four, concept gift for him has the lowest cost per purchase, but its new-customer share is only thirty percent. Concept long-lasting in summer heat costs twenty percent more per purchase, but seventy percent of its buyers are new. Which is better? For an acquisition campaign, the second one, because the first is mostly selling to existing customers who'd probably buy anyway. I check hook rate too: the summer-heat concept stops the scroll well, so its first seconds work. Decision: scale summer heat, keep gift for him in the retargeting-heavy mix, and write the principle, performance claims tied to local climate attract new buyers in Qatar. Next test: same concept, Arabic-first voiceover versus English. One variable, pre-set metric, two weeks.
7:32 Recap and next step
Recap. The algorithm's allocation isn't a test. Combine weekly exploration, controlled A/B tests and lift tests. Design with one hypothesis, one variable, a pre-set metric and decision rule. Read results honestly and write principles. Keep the pipeline full to beat fatigue. Your next step: write a test plan for one creative hypothesis in your business, with all six design decisions filled in.
The problem with "just let the algorithm pick"
Automated campaigns do allocate spend to ads they predict will perform, but that is not a test. Spend allocation is biased toward early winners, ads get different audiences, and the platform's choice reflects its prediction, not a controlled comparison. You need both: let the algorithm scale winners, and run structured tests to learn what makes a winner.
Three testing modes
| Mode | Question | How | When |
|---|---|---|---|
| Exploration in-campaign | Which of these concepts can find an audience? | Add new concepts to an existing Advantage+/PMax/Smart+ campaign; judge after a minimum spend | Always-on, weekly |
| Controlled A/B | Does concept A beat concept B for the same audience? | Platform experiment tools (Meta A/B test, Google Ads experiments, TikTok split test) that split audiences | Big bets: new positioning, offer, format |
| Lift test | Does this creative strategy cause more sales than no ads/other creative? | Conversion lift or geo tests (Module 5) | Strategic shifts, budget decisions |
Designing a controlled creative test
- One hypothesis: "Showing the price upfront increases purchase rate among new customers."
- One variable: price shown vs not shown; everything else equal.
- Primary metric: cost per purchase (or purchase rate) — decided before launch.
- Sample size: enough conversions per cell to detect the effect you care about. As a rough, illustrative planning rule, tests with fewer than ~50–100 conversions per cell rarely give clear answers for moderate effects; use the platform's power estimate where available.
- Duration: at least one full weekly cycle; avoid major sales days unless that is what you are testing.
- Decision rule: e.g., "ship B if it is better with the platform's reported confidence above 80–90% and the effect is at least 10%."
Reading results honestly
- Look at the primary metric first, secondary metrics second.
- Check audience composition — did one ad get more existing customers?
- Beware novelty effects: new ads often spike and fade.
- Record learnings as principles ("problem-first hooks outperform product-first for new customers in KSA"), not just "ad 14 won".
Creative fatigue
Frequency rising, click-through falling, and cost per result creeping up on a stable audience are fatigue signals. Automated campaigns partly manage this by rotating, but a concept eventually saturates its audience. Plan a refresh cadence based on spend: higher spend exhausts concepts faster. A simple rule: always have new concepts in testing so that a replacement is ready before the incumbent fades.
Worked example: a Jeddah food-delivery startup
The startup runs Smart+ on TikTok and Advantage+ on Meta. Each week:
- Monday: 4 new concepts added to each campaign (exploration).
- Thursday: concepts with spend above a threshold and cost per first order below 1.3× target are kept; others paused.
- Monthly: one controlled A/B test on a strategic question (e.g., "Arabic-first voiceover vs English with Arabic captions"). Result: Arabic-first won decisively for first orders; it became a principle.
- Quarterly: a geo lift test on whether TikTok spend is incremental.
Hands-on: a simple concept-level analysis in Python
import pandas as pd
df = pd.read_csv("ad_export.csv") # columns: ad_name, concept, spend, impressions, video_3s, purchases, new_customer_purchases
g = df.groupby("concept").agg(spend=("spend", "sum"), imps=("impressions", "sum"),
v3=("video_3s", "sum"), purch=("purchases", "sum"),
new=("new_customer_purchases", "sum"))
g["cpa"] = g["spend"] / g["purch"].where(g["purch"] > 0)
g["hook_rate"] = g["v3"] / g["imps"]
g["new_share"] = g["new"] / g["purch"].where(g["purch"] > 0)
min_spend = 300 # illustrative threshold in account currency
print(g[g["spend"] >= min_spend].sort_values("cpa").round(3))Name ads with a consistent convention (concept__angle__format__version) so you can group reliably.
A note on statistics without the jargon
Two ads will never perform identically, so the question is always "is this difference bigger than noise?" Platform experiment tools report a confidence or probability figure; treat it as a guide, not a guarantee. Three habits keep you honest: decide the minimum effect worth acting on (a 2% difference is rarely worth a strategy change), do not peek daily and stop the moment one ad looks ahead, and re-test surprising results before rewriting your playbook. If a result contradicts everything you know about your customers, it is more likely noise or a tracking problem than a revelation.
Pitfalls
- Declaring winners on tiny numbers.
- Testing several variables at once in an "A/B".
- Letting the algorithm's spend allocation stand in for a test.
- Forgetting to write down the learning.
How to measure success
Win rate of new concepts, time to find a new winner, cost per result of the top concepts over time, and a growing library of documented creative principles.
Key takeaways
- Platform spend allocation is not a test; combine in-campaign exploration, controlled A/B tests and lift tests.
- Design tests with one hypothesis, one variable, a pre-set metric, adequate sample and a decision rule.
- Read results for audience composition and novelty effects, and record learnings as principles.
- Manage fatigue with a spend-based refresh cadence and a constant testing pipeline.
- Use strict ad naming conventions to analyze at the concept level.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Write a test plan for one creative hypothesis: hypothesis, variable, primary metric, minimum conversions per cell, duration and decision rule.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.