Digital PR, Brand Mentions & E-E-A-TNewsworthy campaign ideation · Lesson 3 of 14
Data-led campaigns: original research, surveys and public datasets
Video lecture
Data-led campaigns: original research, surveys and public datasets
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Data-led campaigns
A journalist gets two hundred pitches a week. Most say some version of our CEO has thoughts. One says: we analyzed official rent and earnings data for twenty cities, here's where young workers spend the biggest share of their pay on rent, and here's your city's figure. Guess which one gets opened. Data gives journalists something they can't easily produce themselves. In this lecture, you'll learn the three sources of campaign data, the survey basics you must get right, how to build a simple index in Python with a proper audit trail, how to use AI safely, and how to turn one dataset into many stories.
0:46 Why data wins
Why do data campaigns earn the strongest coverage? Because editors love a number that changes how readers see their city, their job or their wallet. A good data campaign gives a headline figure, local cuts for regional desks, a methodology journalists can trust, and a link target worth citing. And they age well. An annual index can earn coverage every year, and each edition strengthens your brand's association with the topic. Think of it like a lighthouse. Built once, maintained carefully, and every year, people navigate by it. But a lighthouse in the wrong place, with a faulty lamp, is worse than none at all.
1:32 Three sources of data
Three sources of data. Public data, from national statistics offices, like the UK's Office for National Statistics, the US Census Bureau, the Pakistan Bureau of Statistics, the UAE's Federal Competitiveness and Statistics Center, and Saudi Arabia's General Authority for Statistics. Plus Eurostat, the World Bank, the OECD and open data portals. It's credible and reproducible, but check licenses, definitions and release dates. Surveys, through reputable panel providers, which capture opinions and behaviors you can't observe. And proprietary data, aggregated from your own platform, which is unique, but needs privacy review, consent checks and legal sign-off.
2:13 Survey basics
Survey basics you must get right. First, define the population. UK adults eighteen and over, or UAE residents who shopped online last month. Second, neutral wording. How much do you trust online reviews, not do you agree fake reviews are ruining shopping? Third, sub-group sizes. If you'll report by city or age group, each group needs enough respondents on its own. Fourth, disclosure. Publish the sample size, the population, the dates, the method, and who commissioned it. Research codes, like the UK Market Research Society's Code of Conduct, set expectations for this. Following those norms anywhere builds trust.
2:56 Margin of error
Now the margin of error, which trips up many campaigns. For a simple random sample, the approximate ninety-five percent margin of error for a percentage is one point nine six, times the square root of p times one minus p, divided by n. With a thousand respondents and a fifty percent result, that's about plus or minus three percentage points. So if your survey finds fifty-two percent prefer A and forty-eight percent prefer B, you can't honestly say most people prefer A. Opinion is split. And online panels aren't true random samples, so treat the formula as a rough guide, and never over-claim small differences.
3:42 Worked example 1: rent vs earnings
Worked example one, simple. A recruitment agency in Manchester wants a story on young workers and rent. The team downloads official median rent and median earnings by city from public statistics, checks that both use the same year and definitions, and calculates rent as a share of earnings. The top finding: in some cities, young workers spend a far larger share of pay on rent than the national median. Each city becomes a local story. The page includes the full table, the sources, the download date and a methodology note. Regional desks run the local cuts. Several link to the methodology.
4:26 Worked example 2: Eid wardrobe index (illustrative)
Worked example two, a business scenario with illustrative details. A fashion marketplace in the UAE combines anonymized, aggregated basket data, with legal sign-off, and public inflation data. It estimates the typical cost of an Eid outfit in several Gulf cities, with a clear methodology and the sample period. Regional desks get city cuts. Lifestyle writers get a how to shop smart guide. The landing page offers the full table as a download. And the pre-launch review removes a comparison by nationality that could have stereotyped groups. The index is planned as an annual edition, so each year's coverage builds on the last.
5:10 Watch me do it: pandas index
Watch me do it in Python. I load two CSV files with pandas: median monthly rent by city, and median monthly earnings by city. My load function checks the files exist and have the right columns, so a broken download fails loudly instead of quietly. I merge on city, with a one to one validation, and print any cities that don't appear in both. Then I calculate rent as a percentage of earnings, rounded to one decimal, and rank from most to least affordable. I export the table as a CSV, which becomes the download on the landing page.
5:53 Audit trail + AI rules
Then the part most people skip. I keep the script, the raw files and a short README in one folder. The README lists each source, the download date, the definitions and any exclusions. That folder is my audit trail. When a journalist asks how I calculated a figure, I can answer in minutes. When we publish next year's edition, I rerun the same script. Now, AI. It's great for brainstorming angles, suggesting datasets to investigate, drafting code I review and test, and summarizing results I've already verified. But I never let it provide statistics or quotes. Language models can produce plausible numbers that simply don't exist.
6:39 One dataset, many stories
One dataset, many stories. From a single index you can pitch the national headline, a different lead for each city or region, an angle for business desks on employers, one for personal finance writers with practical tips, and one for trade press in your sector. Each needs its own subject line and first sentence. And plan follow-ups: an update when new official data is released, or a comparison with last year's edition. This is how data campaigns keep earning mentions long after launch day.
7:16 Common mistakes and measures
Common mistakes. Comparing measures with different definitions or years. Reporting tiny sub-groups as if they were reliable. Leading survey questions. Over-claiming differences inside the margin of error. Gating the data behind a form. No methodology, or one written after the analysis. And using AI-generated figures. How do you measure success? Beyond coverage and links, track how often your figures are cited accurately, with the right number and attribution, whether outlets link to the methodology, and whether the index earns repeat coverage in later editions. Those are signs of real authority.
7:55 Recap and try this now
Let's recap. Data campaigns win because they give journalists credible facts, local cuts and a citable link target. Choose public, survey or proprietary data, with the right checks for each. Define your population, word questions neutrally, size sub-groups properly, respect the margin of error, and disclose your method. Keep a reproducible audit trail. And use AI for angles and code, never for statistics. Try this now. Find two public datasets relevant to your brand. Combine them into a simple ranked index, in pandas or a spreadsheet. Write a one-paragraph methodology with sources and dates, and draft three headlines your data genuinely supports.
Why data-led campaigns earn the strongest coverage
Journalists need facts they can't easily produce themselves, and editors love a number that changes how readers see their city, job or wallet. A well-built data campaign gives them a headline figure, local cuts for regional desks, a methodology they can trust and a link target worth citing. Data-led work also ages well: an annual index can earn coverage every year, and each edition strengthens your brand's association with the topic.
The risk is equally clear. A weak sample, a cherry-picked comparison or an AI-invented statistic can end a journalist relationship permanently and may breach advertising and consumer-protection rules. This lesson shows how to do it properly.
Three sources of data
| Source | Examples | Strengths | Watch-outs |
|---|---|---|---|
| Public data | National statistics offices (UK ONS, US Census Bureau and BLS, Pakistan Bureau of Statistics, UAE's Federal Competitiveness and Statistics Center, Saudi Arabia's General Authority for Statistics), Eurostat, World Bank, OECD, open-data portals, regulators | Credible, free, reproducible | Check licenses and definitions; note release dates; don't compare incompatible measures |
| Surveys | Online panels via reputable research providers | Opinions and behaviors you can't observe otherwise | Sample quality, neutral wording, sub-group sizes, disclosure |
| Proprietary data | Aggregated platform data (bookings, searches, prices) | Unique, hard to copy | Privacy, consent, legal review, commercial sensitivity |
Survey basics you must get right
- Define the population ("UK adults 18+", "UAE residents who shopped online in the past month").
- Neutral wording. "How much do you trust online reviews?" not "Do you agree fake reviews are ruining shopping?"
- Sample size for sub-groups. If you'll report by city or age group, each group needs enough respondents on its own; small sub-groups produce unstable percentages.
- Margin of error. For a simple random sample, the approximate 95% margin of error for a percentage is
1.96 × √(p(1−p)/n). With n = 1,000 and p = 50%, that's about ±3.1 percentage points. Online panels are not true random samples, so treat this as a rough guide and don't over-claim small differences. - Disclosure. Publish the sample size, population, dates, method and who commissioned the research. In the UK, members of the British Polling Council and researchers following the Market Research Society Code of Conduct have disclosure obligations; following those norms anywhere builds credibility.
Hands-on: build a city affordability index with Python and pandas
This example combines two public datasets — median rent and median earnings by city — into a ranked index with a reproducible audit trail. Replace the CSV files with real, properly licensed data and cite the sources.
# pip install pandas
import pandas as pd
from pathlib import Path
RENT = Path("data/median_rent_by_city.csv") # columns: city, median_monthly_rent
EARN = Path("data/median_earnings_by_city.csv") # columns: city, median_monthly_earnings
def load(path: Path, cols: list[str]) -> pd.DataFrame:
if not path.exists():
raise FileNotFoundError(f"Missing data file: {path}")
df = pd.read_csv(path)
missing = set(cols) - set(df.columns)
if missing:
raise ValueError(f"{path} is missing columns: {missing}")
return df[cols]
rent = load(RENT, ["city", "median_monthly_rent"])
earn = load(EARN, ["city", "median_monthly_earnings"])
df = rent.merge(earn, on="city", how="inner", validate="one_to_one")
dropped = set(rent.city) ^ set(earn.city)
if dropped:
print("Cities without both measures (excluded):", sorted(dropped))
df["rent_to_income_pct"] = (df.median_monthly_rent / df.median_monthly_earnings * 100).round(1)
df["rank_most_affordable"] = df.rent_to_income_pct.rank(method="min").astype(int)
df = df.sort_values("rank_most_affordable")
df.to_csv("output/affordability_index.csv", index=False) # publish as a download
print(df.head(10).to_string(index=False))
print(f"n cities: {len(df)} | median ratio: {df.rent_to_income_pct.median():.1f}%")Keep the script, the raw files and a short README with sources, download dates and definitions in one folder. That folder is your audit trail when a journalist asks "how did you calculate this?"
Using AI in data campaigns — safely
- Good uses: brainstorming angles, suggesting public datasets to investigate (then verifying them at the source), drafting code you review and test, and writing plain-English summaries of results you've already verified.
- Never: let an AI tool "provide" statistics, estimate figures you present as findings, or generate quotes. Language models can produce plausible numbers that don't exist.
- Label everything. If a figure is an estimate or projection, say so, with the method.
Worked example: a Gulf "cost of the Eid wardrobe" index
A fashion marketplace in the UAE combines anonymized, aggregated basket data (with legal sign-off) and public inflation data to estimate the typical cost of an Eid outfit in several Gulf cities, with a clear methodology and the sample period. Regional desks receive city cuts, lifestyle writers get a "how to shop smart" guide, and the landing page offers the full table as a download. The pre-launch review removes a comparison by nationality that could have stereotyped groups. (Illustrative scenario.)
How to measure success
Beyond coverage and links, track how often your figures are cited accurately (with the correct number and attribution), whether outlets link to the methodology, and whether the index earns repeat coverage in later editions. Those are signs of genuine authority.
Key takeaways
- Data campaigns work because they give journalists credible facts, local cuts and a citable link target.
- Use public, survey or proprietary data — each with its own checks for licenses, sampling, privacy and legal review.
- Define the population, word questions neutrally, size sub-groups adequately and disclose methodology.
- Keep a reproducible audit trail (script, raw data, sources, definitions) for every published figure.
- Use AI for angles, code drafts and summaries — never as a source of statistics.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Find two public datasets relevant to your brand, combine them into a simple ranked index (in pandas or a spreadsheet), write a one-paragraph methodology with sources and dates, and draft three headlines your data genuinely supports.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.