AI for Data Analysis & Decision Making · Statistical traps, experiments and forecasting · lesson 12 of 16 · 12 min
Avoiding common statistical traps
Why this matters more with AI
AI makes it easy to run many analyzes quickly. More analyzes mean more chances to find patterns that are not real, and AI narratives can describe spurious patterns fluently. Knowing the classic traps lets you catch them, whether you or the AI fell in.
Trap 1: correlation is not causation
Two things moving together may share a cause, or be coincidence. Ice cream sales and sunburn rise together because of sunny weather. In business: customers who use the mobile app spend more, but perhaps engaged customers both download the app and spend more. Ask: what else could cause both? Could the arrow run the other way? Only controlled experiments (or careful causal methods) support causal claims.
Trap 2: Simpson's paradox
A trend in combined data can reverse within every subgroup. Illustrative example: a new checkout appears to lower conversion overall, but within both mobile and desktop it improves conversion. The overall drop occurs because the new checkout was released while mobile traffic (which converts lower) grew as a share of visits. Ask: does the pattern hold within key segments? Has the mix changed?
Trap 3: small samples and extreme rates
Small groups produce extreme results by chance. The best and worst performing stores are often the smallest. Ask: how many observations underlie each rate? Would the result survive with more data?
Trap 4: multiple comparisons
Test 20 segments at a conventional significance threshold, and on average about one will look "significant" by chance even if nothing real is happening. AI can test hundreds of cuts in seconds. Ask: how many comparisons were made? Was this hypothesis set before looking? Confirm surprising results with new data.
Trap 5: survivorship bias
Analyzing only what survived distorts conclusions. Studying only your retained customers to learn "what makes customers successful" ignores those who churned doing the same things. Ask: who is missing from this data, and why?
Trap 6: selection bias
Survey respondents, reviewers and beta users are not typical customers. Satisfaction among people who chose to answer is not satisfaction among all customers. Ask: how were these people or records selected?
Trap 7: regression to the mean
Extreme results tend to be followed by more typical ones. The worst month is often followed by a better one, with or without your intervention. If you only act after bad months, your actions will appear to work. Ask: would we have seen improvement anyway?
Trap 8: averages of averages and base effects
Averaging percentages from groups of different sizes gives wrong totals. Large percentage changes on tiny bases ("up 300%!") mislead. Ask: is this weighted correctly? What are the absolute numbers?
Using AI as a trap detector
Here is my analysis and conclusion: [summary].
Act as a skeptical statistician. Check for: correlation vs causation,
Simpson's paradox (mix shifts), small samples, multiple comparisons,
survivorship and selection bias, regression to the mean, and
base-rate issues. For each, say whether it could apply here and what
check would rule it out.
This will not catch everything, but it makes a structured second look fast.
Worked example
A retail chain's AI-assisted analysis finds that stores with the new layout had higher sales growth. Trap check: the new layout was rolled out first to stores that had the worst previous quarter (regression to the mean), and to larger urban stores (selection). Comparing each new-layout store with a similar store without it, over the same period, shows a much smaller effect. The chain decides to run a proper pilot with matched control stores.
Hands-on: see Simpson's paradox in ten lines
Illustrative data, so you can watch the reversal happen:
import pandas as pd
df = pd.DataFrame({
"checkout": ["old", "old", "new", "new"],
"device": ["desktop", "mobile", "desktop", "mobile"],
"sessions": [8000, 2000, 3000, 7000],
"orders": [400, 40, 165, 161],
})
by_device = df.assign(cr=df.orders / df.sessions)
overall = df.groupby("checkout")[["sessions", "orders"]].sum()
overall["cr"] = overall.orders / overall.sessions
print(by_device) # new checkout is better on desktop (5.5% vs 5.0%) AND mobile (2.3% vs 2.0%)
print(overall) # ...but worse overall (3.3% vs 4.4%) because traffic shifted to mobile
The fix is always the same: compare within segments, and ask whether the mix changed.
Hands-on: how many "significant" findings would chance produce?
import numpy as np
rng = np.random.default_rng(7)
# 20 segments where nothing is really different: p-values are uniform under the null
p = rng.uniform(size=20)
print("segments 'significant' at 0.05:", (p < 0.05).sum())
Run it a few times with different seeds. You will regularly see one or two "wins" from pure noise, which is why pre-registered hypotheses and confirmation on new data matter.
Second worked example: a Lahore food-delivery app's "best rider"
An AI summary names the rider with the highest customer rating as a model for training. The trap check shows the rider completed only nine deliveries (small sample) in a quiet neighborhood (selection). Among riders with at least 200 deliveries, ratings cluster tightly, and the differences are within what chance would produce. Training content is instead built from patterns shared by many high-volume riders.
Going further
When decisions are important and you can intervene, prefer experiments: A/B tests online, randomized or matched pilots offline, holdout groups for campaigns. AI can help design them: sample size estimates, randomization plans and analysis scripts. Experiments are the most reliable antidote to most of the traps above.
Video lecture: Avoiding common statistical traps
Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.
- Statistical traps
- Why it matters with AI
- Trap 1: correlation ≠ causation
- Traps 2 and 3
- Traps 4 and 5
- Traps 6 to 8
- Example 1: store layouts
- Example 2: the 'best rider'
- Watch me do it, part 1
- Watch me do it, part 2
- AI as skeptic
- The antidote: experiments
- Measuring progress
- Common mistakes
- Recap and try this now
Lecture transcript
Statistical traps
A retail chain rolls out a new store layout and sales growth in those stores jumps. Success? Not so fast. The new layout went first to the stores that had the worst previous quarter, and to bigger city stores. Both of those facts could explain the jump without the layout doing anything. In this lecture you'll learn the eight statistical traps that AI-assisted analysis falls into most often, how to spot each one, two code demonstrations you can run in minutes, and a prompt that turns AI into a structured skeptic.
Why it matters with AI
Why does this matter more with AI? Because AI makes it easy to run many analyzes quickly. More analyzes mean more chances to find patterns that aren't real, and AI narratives describe spurious patterns just as fluently as real ones. Knowing the classic traps lets you catch them, whether you fell in or the AI did. Think of it like knowing the common scams: once you can name them, you spot them almost instantly.
Trap 1: correlation ≠ causation
Trap one: correlation isn't causation. Two things moving together may share a cause, or be coincidence. Ice cream sales and sunburn rise together because of sunny weather. In business, customers who use the mobile app spend more, but perhaps engaged customers both download the app and spend more. Ask: what else could cause both? Could the arrow run the other way? Only controlled experiments, or careful causal methods, support causal claims.
Traps 2 and 3
Trap two: Simpson's paradox. A trend in combined data can reverse within every subgroup. Illustratively, a new checkout looks worse overall, but within both mobile and desktop it's better. How? It launched while mobile traffic, which converts lower, grew as a share of visits. So ask: does the pattern hold within key segments, and has the mix changed? Trap three: small samples produce extreme results by chance. The best and worst performing stores are often the smallest. Always ask how many observations sit behind each rate.
Traps 4 and 5
Trap four: multiple comparisons. Test twenty segments at the usual five percent significance level, and on average about one will look significant by chance, even if nothing real is happening. AI can test hundreds of cuts in seconds. Ask: how many comparisons were made, and was this hypothesis set before looking? Confirm surprises with new data. Trap five: survivorship bias. Studying only retained customers to learn what makes customers successful ignores those who churned doing the same things. Ask: who's missing from this data, and why?
Traps 6 to 8
Trap six: selection bias. Survey respondents, reviewers and beta users aren't typical customers. Satisfaction among people who chose to answer isn't satisfaction among everyone. Ask how these people or records were selected. Trap seven: regression to the mean. Extreme results tend to be followed by more typical ones. The worst month is often followed by a better one, with or without your intervention. If you only act after bad months, your actions will always appear to work. And trap eight: averages of averages and base effects. Averaging percentages from groups of different sizes gives wrong totals, and up three hundred percent on a tiny base misleads.
Example 1: store layouts
First example, the store layout from the start. The new layout went first to stores with the worst previous quarter, which is regression to the mean, and to larger urban stores, which is selection. Comparing each new-layout store with a similar store without it, over the same period, shows a much smaller effect. The chain decides to run a proper pilot with matched control stores before rolling out further. Two trap checks turned a confident claim into a well-designed test.
Example 2: the 'best rider'
Second example, a business case. A Lahore food-delivery app's AI summary names the rider with the highest customer rating as a model for training. The trap check shows the rider completed only nine deliveries, a small sample, in a quiet neighborhood, which is selection. Among riders with at least two hundred deliveries, ratings cluster tightly, and differences are within what chance would produce. So training content is built from patterns shared by many high-volume riders instead of one lucky outlier.
Watch me do it, part 1
Watch me demonstrate Simpson's paradox with ten lines of pandas and illustrative data. Four rows: old and new checkout, desktop and mobile, with sessions and orders. By device, the new checkout wins on desktop, five and a half percent versus five, and on mobile, two point three versus two. Then I sum by checkout. Overall, the new checkout loses: about three point three percent versus four point four. Why? The new checkout's traffic is mostly mobile. Seeing it happen once makes you look for it forever.
Watch me do it, part 2
Next, multiple comparisons. I generate twenty random p-values for segments where nothing is truly different. Then I count how many are below point zero five. I run it with a few different seeds. Most runs show one or two significant findings from pure noise. Now imagine an AI summary that highlights those one or two, in confident prose, with a chart. That's why hypotheses should be set before looking, and surprises confirmed on new data.
AI as skeptic
Now use AI as a trap detector. Paste your analysis and conclusion and say: act as a skeptical statistician. Check for correlation versus causation, Simpson's paradox and mix shifts, small samples, multiple comparisons, survivorship and selection bias, regression to the mean, and base-rate issues. For each, say whether it could apply here, and what check would rule it out. It won't catch everything, but it makes a structured second look take two minutes.
The antidote: experiments
And when decisions are important and you can intervene, prefer experiments: A B tests online, randomized or matched pilots offline, and holdout groups for campaigns. They're the most reliable antidote to most of these traps, because randomization balances the hidden differences you'd otherwise have to argue about. The next lesson shows how to design and read one properly.
Measuring progress
How do you know your team is getting better at this? Watch what happens to surprising findings. In a healthy team, surprises are labeled as hypotheses, checked against the traps, and confirmed on new data or in a test before they reach a slide. Keep a small log of claims that didn't survive the checks. Over time, that log becomes a brilliant teaching tool, because it shows new analysts exactly which traps your organization falls into most often.
Common mistakes
Common mistakes. Accepting the first plausible explanation. Ignoring group sizes. Testing many cuts and reporting the winners. Studying only survivors. Acting only after extreme months and crediting yourself for the recovery. And averaging percentages across groups of different sizes. Each of these is easy to prevent with the checks you've just seen, and each one, left unchecked, eventually reaches a decision-maker.
Recap and try this now
Recap. More analyzes mean more spurious patterns, so know the traps: correlation versus causation, Simpson's paradox, small samples, multiple comparisons, survivorship and selection bias, regression to the mean, and base effects. Use AI as a structured skeptic, and prefer experiments for important causal questions. Try this now: take one conclusion from a recent analysis, run the skeptical statistician prompt, identify the most plausible trap, and run the check that would rule it out.
Key takeaways
- More analyzes mean more spurious patterns; know the traps.
- Key traps: correlation vs causation, Simpson's paradox, small samples, multiple comparisons, survivorship and selection bias, regression to the mean, base effects.
- Use AI as a structured skeptic to check for each trap.
- Prefer experiments (A/B tests, matched pilots, holdouts) for important causal questions.
Try it
Take one conclusion from a recent analysis and run the skeptical statistician prompt. Identify which trap is most plausible and what check would rule it out.