Digital PR, Brand Mentions & E-E-A-TNewsworthy campaign ideation · Lesson 3 of 14

Data-led campaigns: original research, surveys and public datasets

Article · 8 min · 9 min lecture

Video lecture

Data-led campaigns: original research, surveys and public datasets

12 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 12

Data-led campaigns

  • 200 pitches, one gets opened
  • Three sources of data
  • Survey basics + a Python index
  • AI safely • many stories from one dataset

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why data-led campaigns earn the strongest coverage

Journalists need facts they can't easily produce themselves, and editors love a number that changes how readers see their city, job or wallet. A well-built data campaign gives them a headline figure, local cuts for regional desks, a methodology they can trust and a link target worth citing. Data-led work also ages well: an annual index can earn coverage every year, and each edition strengthens your brand's association with the topic.

The risk is equally clear. A weak sample, a cherry-picked comparison or an AI-invented statistic can end a journalist relationship permanently and may breach advertising and consumer-protection rules. This lesson shows how to do it properly.

Three sources of data

SourceExamplesStrengthsWatch-outs
Public dataNational statistics offices (UK ONS, US Census Bureau and BLS, Pakistan Bureau of Statistics, UAE's Federal Competitiveness and Statistics Center, Saudi Arabia's General Authority for Statistics), Eurostat, World Bank, OECD, open-data portals, regulatorsCredible, free, reproducibleCheck licenses and definitions; note release dates; don't compare incompatible measures
SurveysOnline panels via reputable research providersOpinions and behaviors you can't observe otherwiseSample quality, neutral wording, sub-group sizes, disclosure
Proprietary dataAggregated platform data (bookings, searches, prices)Unique, hard to copyPrivacy, consent, legal review, commercial sensitivity

Survey basics you must get right

  • Define the population ("UK adults 18+", "UAE residents who shopped online in the past month").
  • Neutral wording. "How much do you trust online reviews?" not "Do you agree fake reviews are ruining shopping?"
  • Sample size for sub-groups. If you'll report by city or age group, each group needs enough respondents on its own; small sub-groups produce unstable percentages.
  • Margin of error. For a simple random sample, the approximate 95% margin of error for a percentage is 1.96 × √(p(1−p)/n). With n = 1,000 and p = 50%, that's about ±3.1 percentage points. Online panels are not true random samples, so treat this as a rough guide and don't over-claim small differences.
  • Disclosure. Publish the sample size, population, dates, method and who commissioned the research. In the UK, members of the British Polling Council and researchers following the Market Research Society Code of Conduct have disclosure obligations; following those norms anywhere builds credibility.

Hands-on: build a city affordability index with Python and pandas

This example combines two public datasets — median rent and median earnings by city — into a ranked index with a reproducible audit trail. Replace the CSV files with real, properly licensed data and cite the sources.

# pip install pandas
import pandas as pd
from pathlib import Path

RENT = Path("data/median_rent_by_city.csv")         # columns: city, median_monthly_rent
EARN = Path("data/median_earnings_by_city.csv")     # columns: city, median_monthly_earnings

def load(path: Path, cols: list[str]) -> pd.DataFrame:
    if not path.exists():
        raise FileNotFoundError(f"Missing data file: {path}")
    df = pd.read_csv(path)
    missing = set(cols) - set(df.columns)
    if missing:
        raise ValueError(f"{path} is missing columns: {missing}")
    return df[cols]

rent = load(RENT, ["city", "median_monthly_rent"])
earn = load(EARN, ["city", "median_monthly_earnings"])

df = rent.merge(earn, on="city", how="inner", validate="one_to_one")
dropped = set(rent.city) ^ set(earn.city)
if dropped:
    print("Cities without both measures (excluded):", sorted(dropped))

df["rent_to_income_pct"] = (df.median_monthly_rent / df.median_monthly_earnings * 100).round(1)
df["rank_most_affordable"] = df.rent_to_income_pct.rank(method="min").astype(int)
df = df.sort_values("rank_most_affordable")

df.to_csv("output/affordability_index.csv", index=False)        # publish as a download
print(df.head(10).to_string(index=False))
print(f"n cities: {len(df)} | median ratio: {df.rent_to_income_pct.median():.1f}%")

Keep the script, the raw files and a short README with sources, download dates and definitions in one folder. That folder is your audit trail when a journalist asks "how did you calculate this?"

Using AI in data campaigns — safely

  • Good uses: brainstorming angles, suggesting public datasets to investigate (then verifying them at the source), drafting code you review and test, and writing plain-English summaries of results you've already verified.
  • Never: let an AI tool "provide" statistics, estimate figures you present as findings, or generate quotes. Language models can produce plausible numbers that don't exist.
  • Label everything. If a figure is an estimate or projection, say so, with the method.

Worked example: a Gulf "cost of the Eid wardrobe" index

A fashion marketplace in the UAE combines anonymized, aggregated basket data (with legal sign-off) and public inflation data to estimate the typical cost of an Eid outfit in several Gulf cities, with a clear methodology and the sample period. Regional desks receive city cuts, lifestyle writers get a "how to shop smart" guide, and the landing page offers the full table as a download. The pre-launch review removes a comparison by nationality that could have stereotyped groups. (Illustrative scenario.)

How to measure success

Beyond coverage and links, track how often your figures are cited accurately (with the correct number and attribution), whether outlets link to the methodology, and whether the index earns repeat coverage in later editions. Those are signs of genuine authority.

Key takeaways

  • Data campaigns work because they give journalists credible facts, local cuts and a citable link target.
  • Use public, survey or proprietary data — each with its own checks for licenses, sampling, privacy and legal review.
  • Define the population, word questions neutrally, size sub-groups adequately and disclose methodology.
  • Keep a reproducible audit trail (script, raw data, sources, definitions) for every published figure.
  • Use AI for angles, code drafts and summaries — never as a source of statistics.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A survey of 1,000 adults finds 52% prefer option A and 48% option B. What is the most accurate way to report it?
  2. Which is an appropriate use of an AI assistant in a data-led campaign?
  3. Why keep the script, raw files and a README with sources and definitions in one folder?

Put it into practice

Find two public datasets relevant to your brand, combine them into a simple ranked index (in pandas or a spreadsheet), write a one-paragraph methodology with sources and dates, and draft three headlines your data genuinely supports.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.