AI for Data Analysis & Decision Making · Exploring data and finding insights · lesson 6 of 16 · 8 min
A reusable prompt library for trustworthy AI analysis
Why a prompt library beats clever prompting
Most analysis mistakes with AI come from missing context and missing checks, not from weak wording. A small, shared prompt library fixes both: every prompt carries the same context slots (data grain, definitions, decision) and the same verification requests (show code, counts, limitations). New analysts inherit good habits on day one, and reviewers know what to expect.
The anatomy of a good analysis prompt
| Slot | Purpose | Example | |---|---|---| | Role and mode | Force computation, not guessing | "Use Python to compute every number; show the code." | | Data description | Grain, columns, period, source | "One row per order line, Jan-Sep 2026, from Shopify export." | | Definitions | Remove ambiguity | Paste the glossary section | | Decision | Focus the analysis | "Decide whether to expand next-day delivery to Jeddah." | | Task steps | Make the method explicit | Numbered steps | | Checks | Built-in verification | "Report group sizes; reconcile totals; list 3 limitations." | | Output format | Make review easy | "Table, then chart, then 5-line summary." | | Guardrails | Stop overreach | "No causal claims unless data is from a controlled test; no new numbers in the summary." |
Hands-on: the core library (copy these)
1. Data profile (no changes)
Profile the attached data with Python. Report: rows, columns and types, date range, missing values
per column (count and %), duplicate rows and duplicate IDs, top 20 values of each categorical column,
min/max/median/mean of numeric columns, and suspicious values. Do not modify the data. Show code.
2. Segment comparison with sizes
Compare {metric (definition)} across {segments} for {period}. For every group report n, the metric,
and a 95% confidence interval where it applies. Sort by n. Mark groups with n < {threshold} as
"too small to interpret". Show code. Then list 2 alternative explanations for the largest difference.
3. Driver decomposition
{KPI} changed from {A} to {B} between {period 1} and {period 2}. Decompose the change into
{factors, e.g. traffic x conversion x order value} overall and by {segment}. Report which factor and
which segment explain most of the change, with numbers. Show code and the reconciliation to the total.
4. Skeptical review of my conclusion
Here is my analysis summary and conclusion: {paste}. Act as a skeptical statistician. Check for:
correlation vs causation, Simpson's paradox, small samples, multiple comparisons, selection and
survivorship bias, regression to the mean, and definition problems. For each: could it apply,
and what specific check would rule it out?
5. Executive summary with guardrails
Write a Situation / Insight / Implication / Action summary (max 150 words) for {audience}
using ONLY the numbers in the tables below. Include one sentence on confidence and the main caveat.
No causal language unless the data comes from a controlled test. Plain English.
Versioning and testing prompts
Treat prompts like code. Store them in a shared document or repository with a version number, an owner, and a "known good" example output. When a model changes (vendors update models frequently), rerun the example and compare. If outputs drift (for example, the model stops showing code), update the prompt.
Worked example: a Jeddah logistics team
A logistics company's operations analysts each wrote their own prompts, and results varied wildly: some showed code, some did not; some reported group sizes, most did not. The lead analyst built a five-prompt library with the slots above, added the company glossary, and asked reviewers to reject any analysis that did not follow a library prompt. Within a month, review time dropped because every output had the same shape, and two errors were caught by the built-in reconciliation step before reaching management.
Second worked example: a creator's sponsorship analysis
A US-based YouTube creator uses prompt 2 on a CSV of past sponsored videos (views, watch time, clicks, sponsor category). The assistant reports that one category looks strong but is based on only four videos, marks it "too small to interpret", and suggests two alternative explanations (those videos were also longer and published on weekends). The creator negotiates rates on the categories with enough data and treats the small one as a test.
Chaining prompts into a workflow
The five prompts work best in order: profile (1) before any analysis, segment comparison (2) or decomposition (3) for the question, skeptical review (4) on your draft conclusion, and the guarded summary (5) last, only after numbers are verified. Some teams save this chain as a checklist at the top of every analysis document, with a tick box per step and a link to the verification log.
Pitfalls
- Prompts without definitions or grain, so the model guesses.
- No output format, so every answer looks different and review is slow.
- Treating a prompt as permanent; models change, so rerun known-good examples.
- Asking for a summary before the numbers are verified.
Video lecture: A reusable prompt library for trustworthy AI analysis
Lecture coming soon · 16 chapters · about 8 minutes. Read the full transcript below.
- An analysis prompt library
- Why a library
- Eight slots
- Five core prompts
- The skeptical review
- Prompts are code
- Output format
- Example 1: a creator's sponsorships
- Example 2: Jeddah logistics
- Watch me do it, part 1
- Watch me do it, part 2
- Where it lives
- Common mistakes
- Measuring the library
- Recap
- Try this now
Lecture transcript
An analysis prompt library
Two analysts ask the same AI the same question about the same data. One gets a careful answer with code, group sizes and limitations. The other gets a confident paragraph with a number that's wrong. The difference isn't talent or clever wording. It's what the prompt contained. In this lecture you'll learn the anatomy of a good analysis prompt, a five-prompt library you can copy today, how to version and test prompts when models change, and how a shared library speeds up review.
Why a library
Why a library? Because most analysis mistakes with AI come from missing context and missing checks, not from weak wording. A shared library fixes both. Every prompt carries the same context slots, like data grain, definitions and the decision. And the same verification requests, like show your code, report counts, and list limitations. New analysts inherit good habits on day one, and reviewers know exactly what to expect from every output.
Eight slots
Here's the anatomy, eight slots. Role and mode: use Python to compute every number and show the code. Data description: grain, columns, period and source. Definitions: paste your glossary. Decision: what this informs. Task steps: numbered. Checks: group sizes, reconciliation, limitations. Output format: table, then chart, then a short summary. And guardrails: no causal claims without a controlled test, and no new numbers in the summary. Think of it like a recipe card for a professional kitchen. Every card has the same sections, so any cook can follow any recipe.
Five core prompts
Now the five core prompts. One, data profile: rows, types, date range, missing values, duplicates, top values, numeric ranges, suspicious values, and change nothing. Two, segment comparison with sizes: the metric per segment with n and a confidence interval, sorted by n, small groups marked too small to interpret, and two alternative explanations. Three, driver decomposition: which factor and which segment explain a KPI change, with reconciliation to the total. Four, skeptical review of your conclusion. And five, an executive summary with guardrails.
The skeptical review
Let's look closer at the skeptical review prompt, because it's the one people skip. You paste your summary and conclusion, and ask the model to act as a skeptical statistician. Check for correlation versus causation, Simpson's paradox, small samples, multiple comparisons, selection and survivorship bias, regression to the mean, and definition problems. For each: could it apply, and what specific check would rule it out? It won't catch everything. But it makes a structured second look take two minutes instead of never happening.
Prompts are code
And treat prompts like code. Store them with a version number, an owner and a known-good example output. Vendors update models frequently, and a prompt that reliably produced code and group sizes last quarter might not today. So when a model changes, rerun the example and compare. If the output drifts, for example the model stops showing its code, update the prompt. That's regression testing for prompts, and it takes ten minutes.
Output format
Output format deserves a moment, because it's what makes review fast. Ask for the same shape every time. A data quality note first. Then tables with counts. Then one or two charts. Then a short summary with a confidence line and the main caveat. When every analysis arrives in that order, a reviewer can check it in minutes, skipping straight to the counts and the caveat. It also makes it obvious when something is missing, like a table with rates but no group sizes.
Example 1: a creator's sponsorships
First example, a simple one. A US-based YouTube creator uses the segment comparison prompt on a spreadsheet of past sponsored videos, with views, watch time, clicks and sponsor category. The assistant reports that one category looks strong, but it's based on only four videos, and marks it too small to interpret. It suggests two alternative explanations: those videos were also longer and published on weekends. The creator negotiates rates on the categories with enough data, and treats the small one as a test.
Example 2: Jeddah logistics
Second example, a business case. A Jeddah logistics company's analysts each wrote their own prompts. Some outputs showed code, some didn't. Some reported group sizes, most didn't. The lead analyst built a five-prompt library with the eight slots, added the company glossary, and asked reviewers to reject any analysis that didn't use a library prompt. Within a month, review time dropped because every output had the same shape, and two errors were caught by the built-in reconciliation step before reaching management.
Watch me do it, part 1
Watch me use prompt three, driver decomposition. The KPI: weekly revenue changed from sixty thousand to about forty-eight thousand between two comparable weeks. I ask the assistant to decompose it into sessions, conversion rate and order value, overall and by device, to report which factor and which segment explain most of the change, and to show the code and the reconciliation to the total. It returns a table: conversion rate explains almost all of it, and mobile explains almost all of that.
Watch me do it, part 2
Then I check the reconciliation line: the factor contributions add up to the total change, within rounding. Good. Next I run prompt four on my draft conclusion: the new checkout hurt mobile conversion. The skeptical review flags a possible alternative: a mobile-heavy campaign started the same week, which could change the mix. My check: compare mobile conversion for returning visitors only, who weren't affected by the campaign. It's still down. Now I'm confident enough to recommend a fix, and I've documented why.
Where it lives
Where should the library live? Somewhere your team already works: a shared document, a wiki page, or a small repository if your analysts use code. Each prompt gets a short name, a version, an owner, the slots to fill, and a link to its known-good example. Some teams also turn their most-used prompts into saved instructions or reusable skills inside their AI tool, so the context loads automatically. Whatever you choose, make it the default path, not an optional extra.
Common mistakes
Common mistakes. Prompts without definitions or grain, so the model guesses. No output format, so every answer looks different and review is slow. Treating prompts as permanent when models change. Asking for a summary before the numbers are verified. And collecting fifty prompts nobody maintains. Five good ones, versioned, beat fifty forgotten ones.
Measuring the library
How do you know your library works? Review time falls because outputs share one shape. Errors are caught by built-in checks before stakeholders see them. New analysts produce reviewable work in their first week. And when a model update changes behavior, you notice within days because your known-good examples fail.
Recap
Recap. Missing context and missing checks cause most AI analysis mistakes, so build a library, not clever one-offs. Use eight slots: role and mode, data, definitions, decision, steps, checks, format and guardrails. Start with five prompts: profile, segment comparison, decomposition, skeptical review and guarded summary. And version and re-test them when models change.
Try this now
Try this now. Copy the five prompts from the lesson into a shared document, add your team's glossary, and give each a version number and an owner. Run prompt one and prompt two on a real dataset, save the outputs as known-good examples, and ask a colleague to review one using the output format as a checklist.
Key takeaways
- Most AI analysis mistakes come from missing context and missing checks, not weak wording.
- Good analysis prompts fill eight slots: role/mode, data, definitions, decision, steps, checks, format, guardrails.
- Start with five prompts: profile, segment comparison with sizes, decomposition, skeptical review, guarded summary.
- Version prompts, keep known-good examples, and re-test when models change.
Try it
Create a shared five-prompt library with your glossary, version numbers and owners; run prompts 1 and 2 on real data and save the outputs as known-good examples.