---
title: "Monitoring AI features in production | Optimize All Academy"
description: "Launch is the beginning AI features change behaviour even when your code doesn't: users bring new questions, content changes, providers update models…"
url: https://optimizeall.com/learn/building-ai-products-and-workflows/monitoring-in-production
updated: 2026-10-05
---

Building AI Products & Workflows · Evaluation-driven development and production monitoring · lesson 12 of 18 · 11 min

# Monitoring AI features in production

## Launch is the beginning

AI features change behaviour even when your code doesn't: users bring new questions, content changes, providers update models, and costs drift. Monitoring catches problems before users or finance do.

## What to monitor

**1. Operational health**

- Latency (median and slow tail), error rates, timeouts, rate-limit hits.
- Provider availability and fallback activations.

**2. Cost**

- Tokens and cost per request, per user, per feature and per successful outcome.
- Retry rates and reasoning token usage.
- Budget alerts and anomaly detection.

**3. Quality**

- Automated checks on live outputs (schema validity, groundedness checks, policy checks).
- Sampled human review with rubrics, stratified by topic and risk.
- LLM-judge scoring on samples (calibrated).

**4. User signals**

- Explicit: thumbs up/down, ratings, "report a problem".
- Implicit: edit rate of drafts, acceptance rate of suggestions, regenerate clicks, abandonment, escalation to humans, repeat questions.

**5. Safety and misuse**

- Policy violations, jailbreak or injection attempts, unusual tool usage, data access anomalies.

**6. Drift**

- Changes in input distribution (new topics, languages, lengths).
- Changes in output characteristics (length, refusal rate, abstention rate).
- Changes after model or prompt updates.

## Implicit signals are gold

Explicit feedback is sparse: most users never click thumbs. Implicit signals scale:

- **Edit distance** between AI draft and sent version shows usefulness.
- **Regenerate** clicks suggest dissatisfaction.
- **Escalations** reveal where automation fails.
- **Follow-up questions** like "that's not what I asked" flag misunderstandings.

Define these signals at design time so the product logs them.

## Dashboards and alerts

A practical AI feature dashboard includes: volume, latency, cost per outcome, quality score (sampled), user signals, safety events, and version markers (so you can see changes after deployments). Alerts should fire on: error spikes, cost anomalies, quality drops beyond thresholds, safety events and provider failures. Keep alert thresholds meaningful to avoid alert fatigue.

## Tracing

For complex features (retrieval, chains, agents), log **traces**: each step's inputs, outputs, timings and versions under one request ID. When a user reports a bad answer, you can see whether retrieval, the model, a tool or post-processing caused it. Many observability tools support AI traces; the key is consistent IDs and versions.

## Privacy in monitoring

Monitoring needs data, but logs can accumulate sensitive information. Apply the principles from the privacy lesson: minimise, mask, restrict access, set retention, and document what is logged. Let users know, in your privacy notice, how interactions are reviewed for quality.

## Worked example: a drop nobody noticed

A document summarisation feature's thumbs-down rate is stable, but monitoring shows average output length fell sharply after a provider model update, and the edit rate rose. Sampled review shows summaries now omit key figures. Because prompt, model version and output metrics were logged together, the team quickly links the change to the update, adjusts the prompt to require key figures, reruns evaluations and redeploys. Without the length and edit-rate signals, the degradation would have been noticed only through complaints weeks later.

## Incident response

When monitoring detects a serious issue:

1. Assess severity (harm, scale, data exposure).
2. Mitigate: kill switch, rollback to previous version, fallback mode.
3. Preserve logs and traces.
4. Communicate with affected users and, where required, regulators.
5. Post-incident review: root cause, new evaluation cases, control improvements.

## Who watches the dashboard?

Monitoring only works with ownership. Assign a named owner for each AI feature's dashboard, a weekly review slot to look at sampled outputs and trends, and clear escalation paths. A dashboard nobody looks at is just an expensive log.

## Hands-on: one log event per AI call, then a dashboard query

Monitoring starts with a consistent event. Emit one structured record per model call (and one per tool call for agents), tied together by a request ID and tagged with versions. OpenTelemetry now has semantic conventions for generative AI (attribute names such as `gen_ai.request.model` and `gen_ai.usage.input_tokens`); following them makes your data portable across observability tools. Check the current conventions, which are still evolving.

```python
import json, time, uuid, logging, os
import anthropic

log = logging.getLogger("ai_events")
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
PROMPT_VERSION = "support-draft-v7"

def draft_reply(ticket_text: str, request_id: str | None = None, user_segment: str = "unknown"):
    request_id = request_id or str(uuid.uuid4())
    start = time.monotonic()
    status, usage, output = "ok", None, ""
    try:
        resp = client.messages.create(model=MODEL, max_tokens=1500,
                                      system=open("prompts/support-draft-v7.md").read(),
                                      messages=[{"role": "user", "content": ticket_text}])
        usage = resp.usage
        output = "".join(b.text for b in resp.content if b.type == "text")
        if resp.stop_reason not in ("end_turn", "stop_sequence"):
            status = f"stop_{resp.stop_reason}"
    except anthropic.APIStatusError as e:
        status = f"api_error_{e.status_code}"
    finally:
        log.info(json.dumps({
            "event": "ai_call", "request_id": request_id, "feature": "support_draft",
            "gen_ai.request.model": MODEL, "prompt_version": PROMPT_VERSION,
            "gen_ai.usage.input_tokens": getattr(usage, "input_tokens", None),
            "gen_ai.usage.output_tokens": getattr(usage, "output_tokens", None),
            "cache_read_tokens": getattr(usage, "cache_read_input_tokens", None),
            "latency_ms": round((time.monotonic() - start) * 1000), "status": status,
            "output_chars": len(output), "segment": user_segment,   # no raw text in this log
        }))
    return request_id, output
```

Log user signals (sent unchanged, edited, rejected, regenerated, escalated) as separate events with the same `request_id`. Then a weekly dashboard is a query away:

```sql
SELECT prompt_version,
       COUNT(*)                                         AS calls,
       PERCENTILE_CONT(0.95) WITHIN GROUP (ORDER BY latency_ms) AS p95_latency_ms,
       AVG(CASE WHEN s.signal = 'sent_unchanged' THEN 1.0 ELSE 0 END) AS accept_rate,
       AVG(CASE WHEN s.signal = 'edited' THEN 1.0 ELSE 0 END)         AS edit_rate,
       AVG(output_chars)                                AS avg_output_chars
FROM ai_calls c LEFT JOIN user_signals s USING (request_id)
WHERE c.ts >= CURRENT_DATE - INTERVAL '7 days'
GROUP BY prompt_version ORDER BY prompt_version;
```

Plot these per day with version markers, and alert on sudden changes in edit rate, output length, error rate or cost per call.

## Going further

Run periodic "evaluation replays": re-run a sample of last month's real inputs through the current production version and compare to what users received. This detects silent drift and gives you a continuously refreshed benchmark grounded in real usage.

## Video lecture: Monitoring AI features in production

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Monitoring AI in production
2. Analogy: the uneaten food
3. Six monitoring areas
4. Implicit signals
5. Dashboards, alerts, traces
6. Privacy and ownership
7. Simple example: product descriptions
8. Catching it early + incidents
9. Business example (illustrative)
10. Hands-on in the lesson
11. Common mistakes
12. How you'll know it's good enough
13. Watch me do it: logging + dashboard
14. Recap
15. Try this now (30 minutes)

## Lecture transcript

### Monitoring AI in production

Here's a story that happens more often than anyone admits. An AI summarisation feature has a stable thumbs-down rate, so everyone assumes it's fine. Meanwhile, after a provider model update, summaries got shorter and started dropping key figures. Users quietly fixed them by hand for weeks. Nobody noticed, because nobody was watching the right signals. In this lesson you'll learn what to monitor, why implicit signals are gold, how tracing speeds up diagnosis, and how to respond when something breaks.

### Analogy: the uneaten food

Here's an analogy. Launching an AI feature without monitoring is like launching a new restaurant and never reading the reviews, checking the bins or counting how much food comes back uneaten. Customers rarely complain directly. They just leave food on the plate and don't come back. In AI features, the uneaten food is the edited draft, the regenerated answer and the escalation.

### Six monitoring areas

Launch is the beginning. AI features change even when your code doesn't: users bring new questions, content changes, providers update models and costs drift. So monitor six things. Operational health: latency, errors, timeouts, rate limits, fallbacks. Cost: tokens and spend per request, user, feature and successful outcome. Quality: automated checks on live outputs, sampled human review and calibrated judge scores. User signals. Safety and misuse. And drift in inputs and outputs, especially after model or prompt changes.

### Implicit signals

Explicit feedback is sparse. Most users never click thumbs up or down. Implicit signals scale. Edit distance between the AI draft and what was actually sent shows usefulness. Regenerate clicks suggest dissatisfaction. Escalations to humans show where automation fails. Follow-ups like that's not what I asked reveal misunderstandings. And acceptance rates of suggestions tell you whether the feature earns its place. Define these signals at design time, so the product logs them from day one.

### Dashboards, alerts, traces

A practical dashboard shows volume, latency, cost per outcome, a sampled quality score, user signals, safety events, and version markers so you can see changes after every deployment. Alerts fire on error spikes, cost anomalies, quality drops beyond thresholds, safety events and provider failures. Keep thresholds meaningful, or alert fatigue will train everyone to ignore them. For retrieval, chains and agents, log traces: every step's inputs, outputs, timings and versions under one request ID. When a user reports a bad answer, you can see whether retrieval, the model, a tool or post-processing caused it.

### Privacy and ownership

Monitoring needs data, but logs can quietly accumulate personal information. Apply the privacy lesson: minimise, mask, restrict access, set retention and document what you log. Tell users in your privacy notice how interactions are reviewed for quality. And assign ownership. Every AI feature needs a named dashboard owner, a weekly slot to review sampled outputs and trends, and a clear escalation path. A dashboard nobody looks at is just an expensive log.

### Simple example: product descriptions

A simple example. Your AI writes product descriptions for an online shop. You log each description's length, whether a person edited it before publishing, and how much they changed. One week, average length drops by a third and edits jump. You check the version log: the provider updated the model on Tuesday. You adjust the prompt to specify length and must-include details, re-run your evaluation set, and redeploy. Found in days, not months.

### Catching it early + incidents

Back to our summarisation story, with good monitoring. Output length drops sharply after a provider update, and edit rates rise. Because the prompt version, model and output metrics are logged together, the team links the change to the update within a day, adjusts the prompt to require key figures, re-runs evaluations and redeploys. When something serious happens, follow an incident routine: assess severity, mitigate with a kill switch, rollback or fallback, preserve logs, communicate with affected users and regulators where required, then run a post-incident review that adds new evaluation cases.

### Business example (illustrative)

Illustrative numbers for the summarisation story. Before the update, average summaries were about a hundred and eighty words with a twenty percent edit rate. After, about a hundred and ten words and a forty-five percent edit rate. The alert fired on day two. Two days later, with a prompt fix and a passing evaluation run, length and edits were back to normal. Without monitoring, the first sign would have been a client complaint weeks later.

### Hands-on in the lesson

The hands-on section shows one structured log event per AI call, tied together by a request ID and tagged with model and prompt versions, using attribute names that follow OpenTelemetry's generative AI conventions so your data stays portable. It logs token usage, including cached tokens, latency, status and output length, but no raw text. You'll also get a SQL query for a weekly dashboard showing latency, acceptance rate, edit rate and output length by prompt version.

### Common mistakes

Common mistakes. Monitoring only uptime and errors. Relying on thumbs up and down, which few people click. No version markers on charts, so you can't link changes to deployments. Alerts so noisy that everyone mutes them. Logging raw personal data because it's convenient. And a dashboard with no owner and no weekly review slot.

### How you'll know it's good enough

How will you know your monitoring is good enough? You detect quality drops before users complain. Every chart has version markers. Alerts fire rarely, and when they do, someone acts. A bad answer reported by a user can be traced to its cause in minutes using the request ID. And the weekly review actually happens, with notes on what was found and fixed.

### Watch me do it: logging + dashboard

Watch me do it. I open the logging wrapper. First, I create a request ID and start a timer. Inside the try block I call the model, keep the usage and output, and set the status to the stop reason if it's anything other than a normal end. On an API error, the status records the status code. In the finally block, whatever happened, I log one JSON event: feature, model, prompt version, input, output and cached tokens, latency, status, output length and user segment, and no raw text. Next, the user signals: when an agent sends, edits or rejects the draft, I log a second event with the same request ID. Then the SQL query joins the two, groups by prompt version for the last seven days, and returns calls, slow-tail latency, accept rate, edit rate and average length. I run it and see version seven's edit rate twice version six's, which is my cue to investigate.

### Recap

To recap: monitor operations, cost, quality, user signals, safety and drift, not just uptime. Implicit signals scale better than thumbs. Use version markers, meaningful alerts and end-to-end traces, keep logs privacy-safe, give every dashboard an owner, and rehearse incident response. Your next step is to design a dashboard for one AI feature: two metrics in each of the six areas, and three alerts with thresholds. Next module: unit economics and UX.

### Try this now (30 minutes)

Try this now. For one AI feature, write the structured log event you'd emit for every call: request ID, feature, model, prompt version, tokens, latency, status and output length. Add three user signals you'll log, like edited, regenerated or escalated. Then sketch a dashboard with two metrics per monitoring area and three alerts with thresholds. Give it an owner and a weekly slot.

## Key takeaways

- Monitor operations, cost, quality, user signals, safety and drift, not just uptime.
- Implicit signals (edit rate, regenerations, escalations) scale better than explicit feedback.
- Use dashboards with version markers, meaningful alerts and end-to-end traces with request IDs.
- Apply privacy controls to logs and prepare incident response with kill switches and rollbacks.

## Try it

Design a monitoring dashboard for one AI feature: list two metrics for each of the six categories and three alerts with thresholds.

- [Previous: Evaluation-driven development](https://optimizeall.com/learn/building-ai-products-and-workflows/evaluation-driven-development)
- [Next: Unit economics of AI features](https://optimizeall.com/learn/building-ai-products-and-workflows/unit-economics)
- [All lessons of Building AI Products & Workflows](https://optimizeall.com/learn/building-ai-products-and-workflows)
