---
title: "AI-powered reporting and anomaly detection"
description: "From dashboards to decisions Most marketing reports describe what happened. Useful reports explain why , flag what is unusual , and recommend what to do…"
url: https://optimizeall.com/learn/ai-performance-marketing/ai-powered-reporting
updated: 2026-10-05
---

AI-Powered Performance Marketing · Guardrails and AI-powered reporting · lesson 14 of 15 · 8 min

# AI-powered reporting and anomaly detection

## From dashboards to decisions

Most marketing reports describe what happened. Useful reports explain **why**, flag **what is unusual**, and recommend **what to do next** — with the uncertainty stated. AI can now handle much of the grunt work: pulling data, spotting anomalies, drafting narratives. The analyst's job shifts to defining the right metrics, checking the reasoning and making decisions.

## A reporting architecture that works

1. **Collect**: platform APIs or connectors (Google Ads, Meta Marketing API, TikTok, LinkedIn), GA4 export to BigQuery, backend orders.
2. **Model**: a clean table per day × channel × campaign with spend, platform conversions, backend orders (where attributable), new customers, margin.
3. **Detect**: rules and simple statistics flag anomalies (spend spikes, conversion drops, CPA outliers, tracking breaks).
4. **Explain**: an LLM receives the *numbers and the anomalies* (not raw access to everything) and drafts a narrative with hypotheses.
5. **Review**: a human checks, edits and decides.
6. **Distribute**: weekly memo, Slack summary, dashboard.

Looker Studio, Power BI and similar tools remain the dashboard layer; AI adds the narrative and triage.

## Anomaly detection you can trust

Start with simple, explainable rules before machine learning:

- **Tracking break**: conversions = 0 for a campaign that averaged > 10/day.
- **Spend anomaly**: today's spend > 1.5× the trailing 14-day median for the same weekday.
- **Efficiency drift**: 7-day CPA outside the 28-day mean ± 2 standard deviations.
- **Data gap**: backend orders vs platform conversions ratio changes sharply (often a tag or consent problem).

## Hands-on: anomaly flags plus an LLM narrative

```python
import os, json
import pandas as pd

df = pd.read_csv("daily_channel_metrics.csv", parse_dates=["date"])
# columns: date, channel, spend, conversions, backend_orders
df = df.sort_values("date")
flags = []
for ch, g in df.groupby("channel"):
    g = g.set_index("date")
    last = g.iloc[-1]
    base = g.iloc[-29:-1]
    if base["conversions"].mean() > 10 and last["conversions"] == 0:
        flags.append({"channel": ch, "type": "tracking_break"})
    med = base["spend"].median()
    if med > 0 and last["spend"] > 1.5 * med:
        flags.append({"channel": ch, "type": "spend_spike", "spend": float(last["spend"]), "median": float(med)})
    cpa = (base["spend"] / base["conversions"].where(base["conversions"] > 0)).dropna()
    last_cpa = last["spend"] / last["conversions"] if last["conversions"] else None
    if last_cpa and len(cpa) > 7 and abs(last_cpa - cpa.mean()) > 2 * cpa.std():
        flags.append({"channel": ch, "type": "cpa_outlier", "cpa": round(last_cpa, 2), "mean": round(cpa.mean(), 2)})

summary = df[df["date"] >= df["date"].max() - pd.Timedelta(days=6)].groupby("channel")[["spend", "conversions", "backend_orders"]].sum()
payload = {"week_summary": summary.round(2).to_dict(), "anomalies": flags}

prompt = f'''You are a performance marketing analyst. Using ONLY the data below, write a weekly memo:
1) Three headline findings with numbers. 2) Each anomaly: likely causes (as hypotheses) and the check to confirm.
3) Recommended actions, each with expected impact and risk. If data is insufficient, say so. Do not invent numbers.
DATA: {json.dumps(payload)}'''

# Send `prompt` to your approved LLM provider's API, reading the key from an environment variable, e.g.:
# api_key = os.environ["LLM_API_KEY"]
print(prompt[:500])
```

Keep the model's input to aggregated, non-personal data. Store prompts and outputs so you can audit how a recommendation was made.

## Prompting for honest narratives

- Tell the model to use only provided data and to label hypotheses as hypotheses.
- Ask for the check that would confirm each hypothesis (for example, "compare tag firing in Tag Assistant").
- Require numbers to be quoted from the data; spot-check them.
- Ask for "what would change my mind" — it surfaces uncertainty.

## What a good weekly memo contains

1. **Scoreboard** — spend, backend new customers, blended CAC and margin after ads versus plan and last week.
2. **What changed and why** — two or three findings, each tied to a number and a hypothesis.
3. **Anomalies** — what was flagged, what was checked, what was fixed.
4. **Tests** — status of running experiments and any results with intervals.
5. **Decisions needed** — options, recommendation, risk. Keep it to one page; link to the dashboard for detail.

## Worked example: an agency in Lahore serving UK clients

A 12-person agency reports for 20 UK e-commerce clients. They built a nightly pipeline: connectors into BigQuery, anomaly rules, and an LLM-drafted memo per client. Account managers now spend Monday morning reviewing flagged issues rather than building slides. The first month caught three tracking breaks within a day (previously found after a week). The agency discloses to clients that memos are AI-drafted and human-reviewed.

## Pitfalls

- Letting the LLM compute metrics from raw rows (it may miscalculate); compute in code, narrate with AI.
- Sending personal data or confidential client data to unapproved tools.
- Over-alerting: tune thresholds so people do not ignore alerts.
- Reporting platform ROAS without the backend and incrementality context.

## How to measure success

Time from issue to detection, analyst hours per report, share of recommendations acted on, and decision quality over time (did actions improve outcomes?).

## Video lecture: AI-powered reporting and anomaly detection

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. AI-powered reporting
2. Why it matters
3. Reporting architecture
4. Explainable anomaly rules
5. Golden rule
6. Simple example: Doha fashion client (illustrative)
7. Prompting for honesty
8. Example: Lahore agency, UK clients (illustrative)
9. Pitfalls
10. The one-page memo
11. Measure your reporting
12. Mistakes + try this now
13. Watch me do it: first AI memo (illustrative)
14. Recap and next step

## Lecture transcript

### AI-powered reporting

Picture Monday morning at most agencies. Someone spends three hours exporting numbers into slides that describe last week, and nobody notices the tracking broke on Thursday. In this lecture you'll learn how to build reporting that explains why things happened, flags what's unusual, and recommends what to do, with AI doing the heavy lifting and people doing the thinking.

### Why it matters

Why does this matter? Because the value of a report is the decision it leads to, and most reports lead to none. Here's an analogy. A weather report that says it rained yesterday is history. A forecast that says bring an umbrella tomorrow is useful. And a pilot's instrument panel that beeps when something is wrong is lifesaving. Great marketing reporting is all three: what happened, what to do next, and an alarm when something breaks. AI makes the forecast and the alarm far cheaper to produce, as long as you keep the arithmetic in code and the decisions with people.

### Reporting architecture

Here's the architecture. Collect data from platform APIs or connectors, from GA4's export to BigQuery, and from your backend. Model it into one clean table: day by channel by campaign, with spend, platform conversions, backend orders, new customers and margin. Detect anomalies with rules. Explain them with an LLM that receives only the numbers and the flags. Then a human reviews, edits and decides. Finally, distribute as a memo, a Slack summary or a dashboard.

### Explainable anomaly rules

Start anomaly detection with rules you can explain to a client. A tracking break: conversions drop to zero for a campaign that normally gets more than ten a day. A spend spike: today's spend is more than one and a half times the median for the past two weeks. Efficiency drift: cost per acquisition moves outside two standard deviations of the four-week average. And a data gap: the ratio between backend orders and platform conversions suddenly changes, which usually means a tag or consent problem.

### Golden rule

Now the golden rule of AI reporting: compute in code, narrate with AI. Language models can miscalculate when you hand them raw rows. So your script calculates the weekly summary and the anomaly flags, and only then sends that small, aggregated payload to the model with a clear prompt. The lesson's Python example does exactly that, and reads the API key from an environment variable rather than hard-coding it.

### Simple example: Doha fashion client (illustrative)

Here's a simple worked example. On Tuesday morning, the anomaly script flags that Meta conversions for a Doha fashion client dropped to zero on Monday, while spend and clicks were normal. The LLM memo says: likely a tracking break, not a demand change, and suggests checking the thank-you page tag and any site release on Monday. The account manager checks and finds the developer changed the checkout confirmation page, removing the data layer push. It's fixed by lunchtime. Without the alert, it might have been found at month end, after four weeks of the algorithm learning from zero conversions.

### Prompting for honesty

How do you prompt for an honest narrative? Tell the model to use only the data provided. Ask it to label causes as hypotheses and give the check that would confirm each one, like comparing tag firing in Tag Assistant. Require numbers to be quoted from the data, and spot-check them. And ask what would change its mind. That last question is magic for surfacing uncertainty.

### Example: Lahore agency, UK clients (illustrative)

Here's an illustrative example. A twelve-person agency in Lahore reports for twenty UK e-commerce clients. They built a nightly pipeline: connectors into BigQuery, anomaly rules, and an AI-drafted memo per client. Account managers now spend Monday morning reviewing flagged issues instead of building slides. In the first month they caught three tracking breaks within a day, where previously they'd found them after a week. And they tell clients the memos are AI-drafted and human-reviewed.

### Pitfalls

Watch four pitfalls. Letting the model do arithmetic on raw data. Sending personal or confidential client data to tools your organization hasn't approved. Over-alerting, so people learn to ignore alerts; tune your thresholds. And reporting platform ROAS alone, without backend numbers and incrementality context. Also, keep a log of prompts and outputs, so you can explain how any recommendation was made.

### The one-page memo

What should the weekly memo actually contain? Keep it to one page. Start with the scoreboard: spend, new customers from the backend, blended acquisition cost and margin after ads, compared with plan and last week. Then two or three findings, each tied to a number and a hypothesis. Then the anomalies: what was flagged, what was checked, what was fixed. Then the status of running tests. And finally, decisions needed, with options, a recommendation and the risk. Everything else lives in the dashboard, linked from the memo.

### Measure your reporting

Measure your reporting like a product. How quickly do you detect issues? How many analyst hours does each report take? What share of recommendations get acted on? And over time, did those actions improve outcomes? That's how you prove the pipeline is worth it.

### Mistakes + try this now

Common reporting mistakes. Asking the model to calculate from raw rows. Pasting client data into unapproved tools. Sending so many alerts that everyone ignores them. Writing memos that describe but never recommend. And reporting platform ROAS without the backend scoreboard. Try this now: write one anomaly rule for your account in plain language, including the threshold, like alert if daily conversions fall below forty percent of the fourteen-day median while spend is within twenty percent of normal. Then write who gets the alert and what they check first.

### Watch me do it: first AI memo (illustrative)

Watch me do it. Let's build the first version of an AI-drafted weekly memo for an illustrative Jeddah beauty retailer. Step one, data: a daily table by channel with spend, platform conversions and backend orders, refreshed every morning. Step two, rules: the three from the lesson, tracking break, spend spike, CPA outlier, plus one custom rule, backend orders to platform conversions ratio moving more than thirty percent from its four-week average. Step three, run the script on last week: it flags one spend spike on TikTok, one hundred and eighty percent of the median on Thursday. Step four, prompt: I paste the aggregated payload into our approved LLM with the honesty instructions. The draft says TikTok spend spiked Thursday, possible causes: a budget change or a new campaign; check the change history. It also says Meta's CPA improved, but notes the backend ratio shifted, so improvement might be measurement. Step five, human check: the change history shows a colleague doubled a budget by mistake. Real cause found in two minutes. I edit the memo, add the decision, and send it. Next week the prompt gets one more instruction based on what I had to fix.

### Recap and next step

Recap. Good reports explain why, flag the unusual and recommend actions. Compute in code, narrate with AI, and keep humans deciding. Start with simple, explainable anomaly rules and prompts that demand hypotheses and checks. Your next step: write three anomaly rules for your own accounts with thresholds, and a prompt that turns your weekly numbers into a memo.

## Key takeaways

- Useful reports explain why, flag the unusual and recommend actions with stated uncertainty.
- Compute metrics and anomalies in code; let the LLM narrate aggregated results.
- Start with simple, explainable anomaly rules: tracking breaks, spend spikes, CPA outliers, data gaps.
- Prompt for hypotheses, confirming checks and 'what would change my mind'.
- Keep personal and confidential data out of unapproved tools; log prompts and outputs.

## Try it

Write three anomaly rules for your accounts with thresholds, and a prompt that turns weekly aggregated metrics into a memo with hypotheses and checks.

- [Previous: Guardrails: brand safety, exclusions and policy compliance](https://optimizeall.com/learn/ai-performance-marketing/guardrails-brand-safety-exclusions)
- [Next: Capstone: a full-funnel AI-driven campaign plan with test design](https://optimizeall.com/learn/ai-performance-marketing/capstone-full-funnel-ai-campaign-plan)
- [All lessons of AI-Powered Performance Marketing](https://optimizeall.com/learn/ai-performance-marketing)
