---
title: "Exploratory analysis with AI | Optimize All Academy"
description: "Exploration before conclusions Exploratory data analysis (EDA) means looking at the data broadly before testing specific ideas: distributions, trends…"
url: https://optimizeall.com/learn/ai-for-data-analysis-and-decision-making/exploratory-analysis
updated: 2026-10-05
---

AI for Data Analysis & Decision Making · Exploring data and finding insights · lesson 4 of 16 · 11 min

# Exploratory analysis with AI

## Exploration before conclusions

Exploratory data analysis (EDA) means looking at the data broadly before testing specific ideas: distributions, trends, relationships and anomalies. AI assistants make EDA fast, which is valuable, but speed can tempt you to accept the first interesting pattern. Good EDA is systematic.

## A systematic EDA prompt

```text
Using Python on the attached data (one row per customer subscription):
1. Summarize each numeric column (distribution, median, percentiles).
2. Plot monthly new subscriptions and cancellations over time.
3. Break down churn rate by plan, country and acquisition channel,
   including the number of customers in each group.
4. List the 5 most notable patterns, each with the supporting numbers
   and group sizes.
5. For each pattern, suggest one alternative explanation.
Show all code.
```

Two elements matter most: **group sizes** (a 50% churn rate in a group of 4 customers means little) and **alternative explanations** (built-in skepticism).

## Distributions before averages

Averages hide shape. Revenue per customer is usually highly skewed: a few large customers and many small ones. The mean can be far above what a typical customer spends. Ask for medians, percentiles and histograms alongside means. For skewed metrics, the median or a percentile often communicates "typical" better.

## Trends over time

When exploring trends:

- Check for **seasonality** (monthly or weekly patterns, Ramadan, Black Friday, school holidays).
- Look for **breaks**: tracking changes, pricing changes, product launches, data pipeline issues. A sudden jump is often a measurement change, not a real change.
- Compare **like with like**: year-over-year for seasonal businesses, not just month-over-month.
- Beware **partial periods**: the current month's data is incomplete.

## Segmentation

Breaking data into segments (country, channel, plan, cohort) reveals patterns that totals hide. AI can generate dozens of segment cuts in seconds. The risk is **the garden of forking paths**: look at enough cuts and some will show dramatic differences by chance. Mitigations:

- Decide the most important segments before exploring, based on the business question.
- Report group sizes and treat small groups as anecdotes.
- Treat surprising findings from broad exploration as hypotheses to confirm with new data or a test, not conclusions.

## Cohort analysis

For customer behavior, cohorts (customers grouped by when they joined) are powerful: they separate "our product got better at retaining customers" from "we acquired different customers". AI can produce cohort retention tables quickly:

```text
Build a monthly cohort retention table: rows = signup month, columns =
months since signup, values = % of the cohort still active. Also show the
cohort sizes. Then describe whether recent cohorts retain better or worse,
noting that recent cohorts have fewer months of data.
```

## Worked example: a subscription box

An EDA on a subscription box business shows churn is much higher in the UAE than the UK. Before acting, the analyst asks for group sizes and cohort breakdowns. Findings: the UAE launched recently, so most UAE customers are in their first months, when churn is naturally highest for all countries. Comparing like-for-like cohort months, the difference shrinks considerably. The "UAE problem" was mostly a composition effect. A good EDA prevented a misguided response.

## Asking the AI to challenge you

Useful follow-ups:

- "What else could explain this pattern?"
- "Which of these findings would you be least confident in, and why?"
- "What data would we need to confirm this?"
- "Is there any sign that data quality issues could create this pattern?"

## Failure modes

- Accepting the first dramatic pattern.
- Ignoring group sizes.
- Comparing partial and full periods.
- Mistaking a tracking change for a behavior change.

## Hands-on: a reusable EDA notebook skeleton

Ask the assistant to fill this structure, then save it as your template:

```python
import pandas as pd
import matplotlib.pyplot as plt

df = pd.read_csv("subscriptions.csv", parse_dates=["signup_date", "cancel_date"])

# 1. Quality
print(df.shape, df.dtypes, df.isna().mean().round(3), sep="\n")

# 2. Distributions: median and percentiles beside the mean
print(df["monthly_value"].describe(percentiles=[.1, .25, .5, .75, .9]))

# 3. Trend with partial-period guard
monthly = df.set_index("signup_date").resample("MS").size()
monthly = monthly[monthly.index < pd.Timestamp.today().normalize().replace(day=1)]  # drop current month
monthly.plot(title="New subscriptions per month (complete months only)"); plt.show()

# 4. Segments WITH group sizes
seg = df.groupby("plan").agg(customers=("customer_id", "nunique"),
                             churned=("cancel_date", lambda s: s.notna().sum()))
seg["churn_rate"] = (seg["churned"] / seg["customers"]).round(3)
print(seg.sort_values("customers", ascending=False))

# 5. Cohort retention (share still active N months after signup)
df["cohort"] = df["signup_date"].dt.to_period("M")
```

The comments encode the habits that matter: medians beside means, complete periods only, and group sizes next to every rate.

## Second worked example: a UK retailer's "weekend effect"

An EDA shows weekend orders have a higher average order value. Before recommending weekend promotions, the analyst checks the distribution: a handful of very large B2B orders placed on Saturdays by one trade customer lift the mean. The median weekend order is similar to weekdays. The "weekend effect" becomes a note about one account, not a marketing strategy.

## Going further

Ask the assistant to produce an EDA report as a reproducible notebook with sections for data quality, distributions, trends, segments and open questions. Reviewing a structured notebook is faster and safer than scrolling through a long chat, and it becomes a baseline for the next analysis.

## Video lecture: Exploratory analysis with AI

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

1. Exploratory analysis
2. Why it matters
3. A systematic prompt
4. Distributions first
5. Trend traps
6. Segment discipline
7. Cohorts
8. Anomaly pass
9. Example 1: the weekend effect
10. Example 2: UAE churn
11. Watch me do it, part 1
12. Watch me do it, part 2
13. Three buckets
14. Common mistakes
15. Good exploration
16. Recap and try this now

## Lecture transcript

### Exploratory analysis

Give an AI assistant a dataset and ask what's interesting, and within seconds it will hand you ten dramatic findings. Churn is double in one country. Weekend orders are worth more. Plan B customers stay longer. Some of those will be real. Some will be noise, composition effects, or artifacts of incomplete data. In this lecture you'll learn systematic exploratory analysis: distributions before averages, trends without traps, segmentation with group sizes, cohorts, and how to make the AI challenge its own findings.

### Why it matters

Why does this matter? Because exploration is where hypotheses are born, and bad hypotheses waste months. AI makes exploration fast, which is valuable, but speed tempts you to accept the first interesting pattern. Good exploratory analysis is systematic. It looks broadly, reports group sizes, checks for data breaks and asks what else could explain each pattern. It's the difference between a detective who gathers evidence and one who arrests the first suspect.

### A systematic prompt

Here's a systematic E D A prompt. Using Python on the attached data, one row per customer subscription: summarize each numeric column, with distribution, median and percentiles. Plot monthly new subscriptions and cancellations. Break down churn rate by plan, country and acquisition channel, including the number of customers in each group. List the five most notable patterns, each with supporting numbers and group sizes. And for each pattern, suggest one alternative explanation. Show all code. The two elements that matter most are group sizes and alternative explanations.

### Distributions first

Distributions before averages. Averages hide shape. Revenue per customer is usually highly skewed: a few large customers and many small ones. The mean can sit far above what a typical customer spends. Think of the average wealth in a room when a billionaire walks in. Everyone just got rich on paper, and nobody's life changed. So ask for medians, percentiles and histograms alongside means. For skewed metrics, the median or a percentile usually communicates typical much better.

### Trend traps

Trends without traps. Check for seasonality: weekly and monthly patterns, Ramadan, Black Friday, school holidays. Look for breaks: tracking changes, pricing changes, launches, pipeline issues. A sudden jump is often a measurement change, not a real change. Compare like with like: year over year for seasonal businesses. And beware partial periods. The current month is incomplete, so a chart that includes it will almost always show a scary drop at the end.

### Segment discipline

Segmentation with discipline. Breaking data into segments reveals patterns totals hide. But AI can generate dozens of cuts in seconds, and look at enough cuts and some will show dramatic differences by chance. Statisticians call it the garden of forking paths. So decide the most important segments before exploring, based on the business question. Report group sizes and treat small groups as anecdotes. And treat surprising findings from broad exploration as hypotheses to confirm with new data or a test, not conclusions.

### Cohorts

Cohorts are powerful for customer behavior. Grouping customers by when they joined separates two stories that totals mix together: our product got better at retaining customers, versus we acquired different customers. Ask the AI for a monthly cohort table: rows are signup months, columns are months since signup, values are the share still active, with cohort sizes. Then ask whether recent cohorts retain better or worse, noting that recent cohorts have fewer months of data.

### Anomaly pass

And anomalies deserve their own pass. Ask the AI to list the ten most unusual days, customers or orders, with the full rows, and a guess at why each is unusual. Some will be data errors, like duplicated imports or test transactions. Some will be real events, like a viral post or a bulk order. Some will be your most important customers. Don't delete them until you know which is which. Anomalies are often where the most valuable insight hides.

### Example 1: the weekend effect

First example, a simple one. An exploratory analysis shows a UK retailer's weekend orders have a higher average order value. Before recommending weekend promotions, the analyst checks the distribution. A handful of very large business orders placed on Saturdays by one trade customer lift the mean. The median weekend order is similar to weekdays. The weekend effect becomes a note about one account, not a marketing strategy.

### Example 2: UAE churn

Second example, a business case. An analysis of a subscription box business shows churn much higher in the UAE than the UK. Before acting, the analyst asks for group sizes and cohorts. The UAE launched recently, so most UAE customers are in their first months, when churn is naturally highest everywhere. Comparing like-for-like cohort months, the difference shrinks considerably. The UAE problem was mostly a composition effect. Good exploration prevented a misguided response, like cutting UAE marketing.

### Watch me do it, part 1

Watch me run the notebook skeleton from the lesson. Section one prints shape, types and the share of missing values per column. Section two prints the distribution of monthly value with percentiles from the tenth to the ninetieth. Section three resamples signups by month and drops the current, partial month before plotting. Section four groups by plan with customer counts, churned counts and churn rate, sorted by group size. And section five sets up cohorts by signup month. Each section encodes a habit.

### Watch me do it, part 2

Then I ask the AI to challenge its own findings. What else could explain this pattern? Which of these findings would you be least confident in, and why? What data would we need to confirm this? Is there any sign that data quality issues could create this pattern? For the plan B finding, it points out that plan B was only offered to annual subscribers, so plan and commitment are tangled. That single question saved me from a wrong conclusion.

### Three buckets

Before you share exploration results, sort them into three buckets. Findings: patterns with enough data, checked group sizes, and no obvious alternative explanation. Hypotheses: interesting patterns that need confirmation with new data or a test. And notes: anomalies or data quality issues to fix. This simple labeling stops a hypothesis from being quoted as a fact in next week's meeting, which is how most analysis goes wrong in organizations.

### Common mistakes

Common mistakes. Accepting the first dramatic pattern. Ignoring group sizes. Comparing partial and full periods. Mistaking a tracking change for a behavior change. And reporting every cut the AI produced instead of the ones tied to the question. Each of these is easy to prevent with the checks you've just seen, and each one, left unchecked, eventually reaches a decision-maker.

### Good exploration

How do you measure good exploration? Every reported pattern comes with a group size and at least one alternative explanation. Surprising findings are labeled as hypotheses with a plan to confirm them. And your exploration notebook is saved as a template, so the next analysis starts from good habits rather than a blank chat.

### Recap and try this now

Recap. Explore systematically: distributions, trends, segments and anomalies, with code shown. Use medians and percentiles for skewed metrics. Watch seasonality, breaks and partial periods. Report group sizes and treat broad-exploration surprises as hypotheses. Use cohorts to separate composition from behavior, and make the AI challenge its own findings. Try this now: run a systematic exploration on a dataset you know, using the prompt or notebook from the lesson. For the most surprising finding, check the group sizes, list two alternative explanations, ask the four challenge questions, and decide whether it's a finding or a hypothesis.

## Key takeaways

- Explore systematically: distributions, trends, segments and anomalies, with code shown.
- Always report group sizes and ask for alternative explanations.
- Use medians and percentiles for skewed metrics; check seasonality, breaks and partial periods in trends.
- Treat patterns found by broad exploration as hypotheses; cohorts help separate composition from behavior.

## Try it

Run a systematic EDA prompt on a dataset you know. For the most surprising finding, check group sizes and list two alternative explanations.

- [Previous: Cleaning and preparing messy data with AI](https://optimizeall.com/learn/ai-for-data-analysis-and-decision-making/cleaning-data-with-ai)
- [Next: Asking better analytical questions](https://optimizeall.com/learn/ai-for-data-analysis-and-decision-making/asking-analytical-questions)
- [All lessons of AI for Data Analysis & Decision Making](https://optimizeall.com/learn/ai-for-data-analysis-and-decision-making)
