Skip to content

Email Marketing & Automation · Testing, metrics, reporting and compliance · lesson 12 of 16 · 15 min

A/B testing and email metrics

The metrics that matter

| Metric | Formula | Notes | |---|---|---| | Delivery rate | Delivered ÷ Sent | Low rates signal bad data or blocks | | Bounce rate | Bounces ÷ Sent | Separate hard and soft | | Open rate | Unique opens ÷ Delivered | Inflated by Apple Mail Privacy Protection; use directionally only | | Click rate (CTR) | Unique clicks ÷ Delivered | A more reliable engagement measure | | Click-to-open rate (CTOR) | Unique clicks ÷ Unique opens | Content effectiveness among openers; affected by open inflation | | Conversion rate | Conversions ÷ Delivered (or ÷ Clicks – state which) | Ties email to goals | | Revenue per recipient (RPR) | Revenue ÷ Delivered | Excellent for comparing campaigns and flows | | Unsubscribe rate | Unsubscribes ÷ Delivered | Rising rate = relevance or frequency issues | | Spam complaint rate | Complaints ÷ Delivered | Keep below 0.1% (Gmail's guidance); never approach 0.3% | | List growth rate | (New subscribers − Unsubscribes − Bounces − Complaints) ÷ List size | Is the list healthy and growing? |

Illustrative calculation

A campaign is sent to 20,000 people; 19,600 are delivered; 588 people click; 49 buy for a total of PKR 245,000.

  • Delivery rate = 19,600 ÷ 20,000 = 98%
  • Click rate = 588 ÷ 19,600 = 3%
  • Conversion rate (of delivered) = 49 ÷ 19,600 = 0.25%
  • Revenue per recipient = 245,000 ÷ 19,600 = PKR 12.5

A/B testing basics

An A/B test (split test) sends two versions to random subsets of your audience and compares a chosen metric.

What to test (in rough order of impact):

  1. Offer or content angle – free delivery vs a discount; tips vs story.
  2. Subject line and preview text – judge by clicks or revenue rather than opens.
  3. CTA – wording, placement, button vs text link.
  4. Design – plain-text style vs designed template.
  5. Send time and day.
  6. Flow timing – 1 hour vs 4 hours for the first cart reminder.

How to run a sound test:

  • One variable at a time.
  • Decide the success metric in advance (clicks, conversions or revenue per recipient).
  • Use a large enough sample. On small lists, differences of a few clicks are often noise. Many platforms offer a winner-selection feature that tests on a portion of the list and sends the winner to the rest; make sure the test portion is large enough.
  • Allow enough time for results before choosing a winner (clicks and purchases can arrive over hours or days).
  • Record results and learnings in a test log.
  • Re-test important findings – what worked last year may not work now.

Beyond A/B: holdout groups

To measure the true impact of a flow or campaign, hold back a small random group that does not receive it and compare their purchase behaviour with those who did. This shows incremental revenue rather than revenue that might have happened anyway.

Benchmarks

Published email benchmarks vary widely by industry, region, list size and how opens are counted. Use them cautiously; your own trends over time are more meaningful.

Reporting email performance

A monthly email report should cover:

  • List growth and health (bounces, complaints, unsubscribes).
  • Campaign performance (click rate, conversions, RPR) with the top and bottom performers.
  • Flow performance (revenue or conversions per flow).
  • Tests run and learnings.
  • Next month's plan.

Common mistakes

  • Choosing subject line winners based on opens alone.
  • Declaring winners from tiny samples.
  • Testing several variables at once.
  • Ignoring list health metrics while celebrating revenue.

Hands-on: how big does a test need to be?

A common rule of thumb for comparing two rates (roughly 80% power at a 5% significance level) is:

Recipients needed PER VARIANT ≈ 16 × p × (1 − p) ÷ d²
  p = your current rate (for example click rate 0.03)
  d = the smallest absolute difference you care about (for example 0.006 = 0.6 points)

In Google Sheets:  =ROUNDUP(16*B1*(1-B1)/B2^2, 0)
Example: p = 0.03, d = 0.006  ->  about 12,900 recipients per variant

If your list is smaller, test bigger changes (offer or angle rather than a word in the subject), run the same test across several sends and combine results, or accept the result as directional.

Hands-on: measuring a flow with a holdout

Holdout design: 10% of new subscribers randomly receive NO welcome flow for 30 days.
Revenue per person (flow group)     = revenue from flow group ÷ people in flow group
Revenue per person (holdout group)  = revenue from holdout ÷ people in holdout
Incremental revenue per person      = flow RPP − holdout RPP
Incremental revenue (total)         = incremental RPP × people who received the flow

Illustrative: flow group PKR 310 per person, holdout PKR 190 → incremental PKR 120 per person. With 8,000 subscribers entering the flow in a quarter, that is about PKR 960,000 of genuinely extra revenue – a far more honest number than "all revenue from anyone who received a flow email".

Worked example 2: a Dubai course creator's subject-line test

A course creator with 6,000 subscribers tests two subject lines. Version A: "Your first client in 30 days" (148 clicks from 3,000). Version B: "The pricing mistake I made for two years" (171 clicks from 3,000). The difference looks meaningful, but with rates around 5% and 3,000 per variant, the sample-size rule says a difference this small is hard to confirm in one send. She repeats the same angle comparison in the next two newsletters; the story-led angle wins all three times. Now she trusts it and briefs more story-led subjects.

How to measure success

  • Every test has a written hypothesis, primary metric, sample size check and stop date.
  • A test log with results, confidence notes and what changed as a result.
  • At least one holdout per quarter on a major flow, reported as incremental revenue.

Video lecture: A/B testing and email metrics: numbers you can trust

Lecture coming soon · 11 chapters · about 8 minutes. Read the full transcript below.

  1. Metrics and testing
  2. Core formulas
  3. Example 1: campaign maths (illustrative)
  4. Why RPR
  5. A/B testing
  6. Sample size rule of thumb
  7. Example 2: Dubai creator (illustrative)
  8. Holdout groups
  9. Benchmarks, reports, mistakes
  10. Testing programme success
  11. Recap and try this now

Lecture transcript

Metrics and testing

Here's a scene from many marketing meetings. Someone says: subject line B won, it had a higher open rate. Everyone nods. But B made less money, the sample was tiny, and half the opens came from privacy features rather than people. Email is full of numbers, and some of them mislead. In this lecture you'll learn which email metrics matter and their exact formulas, how to calculate them from a real campaign, how to run A/B tests that give trustworthy answers, including how big a test needs to be, and how to use holdout groups to measure what email really adds.

Core formulas

Let's start with the metrics and their formulas. Delivery rate: delivered divided by sent. Bounce rate: bounces divided by sent, separating hard and soft. Open rate: unique opens divided by delivered, but inflated by Apple's privacy protection, so directional only. Click rate: unique clicks divided by delivered, far more reliable. Click-to-open rate: clicks divided by opens, which is affected by open inflation too. Conversion rate: conversions divided by delivered, or by clicks, so state which. Revenue per recipient: revenue divided by delivered. Unsubscribe rate and spam complaint rate, each divided by delivered. And list growth rate.

Example 1: campaign maths (illustrative)

Let's calculate from a real campaign, with illustrative numbers. A campaign is sent to twenty thousand people. Nineteen thousand six hundred are delivered. Five hundred and eighty-eight people click. Forty-nine buy, for a total of two hundred and forty-five thousand rupees. Delivery rate: nineteen thousand six hundred divided by twenty thousand, ninety-eight percent. Click rate: five hundred and eighty-eight divided by nineteen thousand six hundred, three percent. Conversion rate of delivered: forty-nine divided by nineteen thousand six hundred, zero point two five percent. And revenue per recipient: two hundred and forty-five thousand divided by nineteen thousand six hundred, twelve and a half rupees.

Why RPR

Why is revenue per recipient so useful? Because it lets you compare any two emails fairly: a big newsletter and a small segment send, a campaign and a flow, this month and last month. It combines delivery, engagement and conversion into one number that matters to the business. Think of it like revenue per visitor in a shop. A busier day isn't necessarily a better day. What matters is how much each visitor was worth. For flows, use revenue per person who entered the flow. And always pair it with the health metrics: unsubscribes and complaints. Revenue that burns your list isn't a win.

A/B testing

Now A/B testing. An A/B test sends two versions to random subsets and compares one chosen metric. What should you test, roughly in order of impact? The offer or content angle first. Then subject line and preview, judged on clicks or revenue, not opens. Then the call to action. Then design, like plain text versus a designed template. Then send time. And flow timing, like one hour versus four hours for the first cart reminder. Rules for a sound test: one variable at a time, the success metric decided in advance, enough sample, enough time for clicks and purchases to arrive, and results written in a test log.

Sample size rule of thumb

So how big does a test need to be? Here's a useful rule of thumb from the lesson text. The recipients you need per variant are roughly sixteen, times your current rate, times one minus that rate, divided by the square of the smallest difference you care about. For example, with a three percent click rate, if you want to detect a difference of zero point six percentage points, you need about twelve thousand nine hundred recipients per version. That surprises most people. On a smaller list, test bigger changes, like the offer or angle, repeat the test across several sends, or treat the result as directional.

Example 2: Dubai creator (illustrative)

Let's do a realistic example. A course creator in Dubai with six thousand subscribers tests two subject lines. A: your first client in thirty days, with a hundred and forty-eight clicks from three thousand. B: the pricing mistake I made for two years, with a hundred and seventy-one clicks from three thousand. B looks better. But with click rates around five percent and three thousand per version, a difference that size is hard to confirm in a single send. So she repeats the same angle comparison, practical promise versus personal story, in her next two newsletters. The story-led angle wins all three times. Now she trusts it.

Holdout groups

Beyond A/B tests, there's an even more honest measure: the holdout group. Hold back a small random group, say ten percent of new subscribers, who don't receive a flow for thirty days, and compare their revenue per person with the group that did. The difference is the incremental revenue, what the flow truly added. For example, with illustrative numbers, the flow group earns three hundred and ten rupees per person, the holdout one hundred and ninety. That's a hundred and twenty rupees of genuine extra revenue per person. Multiply by eight thousand people entering the flow in a quarter, and you get about nine hundred and sixty thousand rupees. That's a number you can defend.

Benchmarks, reports, mistakes

A word on benchmarks. Published email benchmarks vary widely by industry, region, list size and how opens are counted, so use them cautiously. Your own trends over time are far more meaningful. A good monthly email report covers list growth and health: bounces, complaints and unsubscribes. Campaign performance: click rate, conversions and revenue per recipient, with top and bottom performers. Flow performance, ideally with holdout results. Tests run and what you learned. And next month's plan. Common mistakes: picking subject-line winners on opens, declaring winners from tiny samples, testing several variables at once, and celebrating revenue while list health quietly declines.

Testing programme success

How do you measure success of your testing programme itself? Every test has a written hypothesis, a primary metric, a sample size check and a stop date. You keep a test log with results, confidence notes, and what changed as a result, because a test that doesn't change anything was just a curiosity. And you run at least one holdout per quarter on a major flow, reported as incremental revenue. Over a year, that discipline turns your email programme from guesswork into a system that gets measurably better every quarter.

Recap and try this now

Let's recap. Prioritise click rate, conversions and revenue per recipient, and treat opens as directional. Monitor list health: bounces, unsubscribes, complaints and growth. Test one variable at a time with a metric decided in advance, check your sample size with the rule of thumb, repeat tests on smaller lists, and log what you learn. Use holdout groups to measure the incremental value of your flows. Here's your try this now. Calculate delivery rate, click rate, conversion rate and revenue per recipient for your last three campaigns, then use the sample-size formula in the lesson text to plan one A/B test your list can actually support.

Key takeaways

  • Prioritise click rate, conversions and revenue per recipient; treat opens as directional because of privacy features.
  • Monitor list health: bounces, unsubscribes, complaints and list growth.
  • A/B test one variable at a time with a pre-set metric, sufficient sample and time, and log learnings.
  • Use holdout groups to measure incremental impact of flows and campaigns.

Try it

Calculate delivery rate, click rate, conversion rate and revenue per recipient for your last three campaigns (or the illustrative numbers). Plan one A/B test with a hypothesis, metric and sample size.