---
title: "Online monitoring, SLOs and drift | Optimize All Academy"
description: "Offline evals are not enough Offline evals test the cases you thought of. Production brings new users, new phrasing, new documents, provider changes and…"
url: https://optimizeall.com/learn/llm-evals-and-observability/online-monitoring-and-slos
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Production observability, monitoring and experimentation · lesson 12 of 16 · 14 min

# Online monitoring, SLOs and drift

## Offline evals are not enough

Offline evals test the cases you thought of. Production brings new users, new phrasing, new documents, provider changes and traffic spikes. **Online monitoring** watches quality, safety, latency and cost continuously on real traffic, and turns surprises into alerts and new test cases.

## Service level objectives for LLM features

Borrow SRE practice: define **SLIs** (indicators) and **SLOs** (targets) that reflect user experience.

| SLI | Example SLO (illustrative; set your own) |
|---|---|
| Time to first token (streaming) | p95 below an agreed threshold during business hours |
| End-to-end latency per request | p95 and p99 targets per route |
| Error rate (provider errors, timeouts, tool failures) | Below an agreed percentage over 28 days |
| Cost per conversation or per resolved ticket | Below budget, with alert at 80% of budget |
| Online quality pass rate (sampled judges) | At or above offline baseline minus a tolerance |
| Safety violations | Zero on blocking categories; any occurrence pages on-call |

Latency for LLMs depends on input and output length, so track **tokens per request** alongside latency, and consider separate SLOs for short and long tasks. Use percentiles, not averages.

## Cost monitoring

- Compute cost from token attributes and your current price sheet (keep prices in config, not code; they change).
- Break down by route, customer tier, model and prompt version.
- Alert on **cost per successful outcome**, not just total spend: a cheaper model that fails more may cost more per resolved ticket.
- Watch for runaway loops in agents (step counts, repeated tool calls) and set hard limits.

## Online quality evaluation

You cannot label all production traffic, but you can:

1. **Run code checks on 100% of traffic:** format validity, forbidden patterns, PII detectors, tool-call rules, refusal markers, length limits.
2. **Run LLM judges on a sample** (for example a few percent, or stratified by route and risk), using the same calibrated judges as offline.
3. **Use implicit signals:** thumbs up/down, escalation to human, user rephrasing the same question, conversation abandonment, repeated contacts within 24 hours.
4. **Compare with the offline baseline**; alert on sustained drops, not single bad answers.

## Detecting drift

Things drift even when you change nothing:

- **Input drift:** new topics, languages, seasonal questions (Ramadan and Eid shopping, back-to-school, Black Friday, Diwali or White Friday sales in the Gulf). Monitor topic clusters and language mix.
- **Knowledge drift:** documents change; retrieval indexes go stale.
- **Provider drift:** model aliases updated, behavior tweaks, new safety filters. Pin versions where possible and watch finish reasons, refusal rates and output length.

## Hands-on: an online monitor job

```python
# online_monitor.py: hourly job over recent traces exported as JSONL (adapt to your backend's API)
import json, random, statistics
from evals.graders import contains_none, EMAIL
from evals.judge import judge

SAMPLE_RATE = 0.05
traces = [json.loads(l) for l in open("last_hour_traces.jsonl", encoding="utf-8")]

pii_hits = [t["id"] for t in traces if contains_none(t["output"], [EMAIL])[0] == 0.0]
latencies = sorted(t["latency_ms"] for t in traces)
p95 = latencies[int(0.95 * (len(latencies) - 1))] if latencies else 0
cost = sum(t["cost_usd"] for t in traces)
resolved = sum(1 for t in traces if t.get("outcome") == "resolved") or 1

sample = [t for t in traces if random.random() < SAMPLE_RATE and t.get("policy")]
verdicts = [judge(t["policy"], t["country"], t["input"], t["output"])["verdict"] for t in sample]
pass_rate = verdicts.count("PASS") / max(1, len([v for v in verdicts if v != "ERROR"]))

report = {"n": len(traces), "pii_hits": pii_hits, "p95_ms": p95,
          "cost_per_resolved": cost / resolved, "judge_pass_rate": pass_rate, "judged": len(sample)}
print(json.dumps(report))
# send to your alerting system; page on any PII hit, warn on pass_rate below baseline - tolerance
```

Failures found this way should be added to the offline dataset after review.

## Worked example

A UK retailer's assistant held steady for months, then online judge pass rate for delivery questions dropped over a week. Offline evals still passed. Trace slicing showed a surge of questions about a new "nominated day" delivery option that did not exist in the knowledge base yet. The assistant was guessing. The fix was content (add the policy), plus 20 new eval cases and an abstention check for unknown delivery options.

## Pitfalls

- **Averages instead of percentiles.**
- **Alerting on single bad answers** (alert fatigue) instead of sustained trends and blocking categories.
- **Judging production with uncalibrated judges.**
- **Hard-coded prices** that silently go stale.

## How to measure success

SLOs defined and visible per route, alerts that fire rarely and meaningfully, and a steady flow of production failures into offline datasets.

## Video lecture: Online monitoring, SLOs and drift

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Online monitoring, SLOs and drift
2. Analogy: crash test vs dashboard
3. SLIs → SLOs
4. Latency details
5. Cost monitoring
6. Online quality
7. Drift
8. Hands-on monitor job
9. Case: UK retailer delivery drift
10. Example: an SLO for order status
11. Scenario: cost per resolution triples (illustrative)
12. Common mistakes
13. Deeper: two-tier quality alerts
14. Watch me do it: the hourly monitor (illustrative)
15. Recap

## Lecture transcript

### Online monitoring, SLOs and drift

Your offline evals passed. You shipped. Three weeks later, quality quietly slides, costs creep up, and nobody notices until a customer posts a screenshot. Offline evals test the cases you imagined. Production brings everything else. In this lecture you will learn to set service level objectives for LLM features, monitor cost and quality online, and detect drift before your users do.

### Analogy: crash test vs dashboard

An analogy for online monitoring. A car's crash test tells you it was safe on the day it was tested. The dashboard warning lights tell you it is safe today, on this road, in this weather. Offline evals are your crash test. Online monitoring is your dashboard: engine temperature, fuel, tire pressure, all watched continuously, with a warning light before something breaks.

### SLIs → SLOs

Borrow a practice from site reliability engineering. A service level indicator is something you measure that reflects user experience. A service level objective is the target. For LLM features, useful indicators include time to first token when streaming, end-to-end latency per route, error rate across provider errors, timeouts and tool failures, cost per conversation or per resolved ticket, the online quality pass rate, and safety violations, where any occurrence in a blocking category should page someone.

### Latency details

Two details matter for latency. Use percentiles like p95 and p99, never averages, because averages hide the slow experiences users remember. And remember that LLM latency depends heavily on input and output length, so track tokens per request alongside latency, and consider separate objectives for short questions and long tasks.

### Cost monitoring

Cost next. Compute it from token counts and your current price sheet, stored in configuration because prices change. Break it down by route, customer tier, model and prompt version. And alert on cost per successful outcome, not just total spend. A cheaper model that fails more often can cost more per resolved ticket once you count escalations. For agents, also watch step counts and repeated tool calls, and set hard limits against runaway loops.

### Online quality

For quality, you cannot label all production traffic, but you can do a lot. Run code checks on every single response: format, forbidden patterns, PII detection, tool-call rules and length limits. Run your calibrated judges on a sample, perhaps a few percent, stratified by route and risk. Use implicit signals like thumbs down, escalations, users rephrasing the same question, abandonment and repeat contacts. Then compare with your offline baseline and alert on sustained drops, not single bad answers.

### Drift

Things drift even when you change nothing. Input drift: new topics, new languages, seasonal spikes like Ramadan and Eid shopping, back to school, or White Friday sales in the Gulf. Knowledge drift: documents change and indexes go stale. And provider drift: model aliases update, behavior shifts, safety filters change. Pin model versions where you can, and watch finish reasons, refusal rates and output length as early warning signs.

### Hands-on monitor job

The lesson includes an hourly monitoring job in Python. It reads the last hour of traces, runs a PII check on every output, computes p95 latency and cost per resolved conversation, samples five percent for the policy judge, and prints a report you can send to your alerting system. Page on any PII hit, and warn when the judge pass rate falls below baseline minus your tolerance. Every failure it finds should be reviewed and added to your offline dataset.

### Case: UK retailer delivery drift

A worked example. A UK retailer's assistant was stable for months. Then the online judge pass rate for delivery questions dropped over a week, while offline evals still passed. Slicing traces showed a surge of questions about a new nominated-day delivery option that was not in the knowledge base yet. The assistant was guessing. The fix was content, twenty new eval cases, and an abstention check for unknown delivery options.

### Example: an SLO for order status

A simple example of an SLO. Route: order status questions. Indicator: time to first token. Objective: ninety-five percent of requests under an agreed threshold during business hours, measured over twenty-eight days. Error budget: the five percent that may be slower. If a new retrieval change burns half the monthly budget in two days, the alert fires, and the team rolls back before customers notice a pattern. The objective turns vague complaints about slowness into a clear, measurable promise.

### Scenario: cost per resolution triples (illustrative)

Now a realistic scenario with illustrative numbers. A marketplace in Karachi tracks cost per resolved conversation for its seller-support assistant. For months it sits steady. Then, over one weekend, it nearly triples, while total traffic barely moves. The breakdown by route shows the cause: a new onboarding flow sends long policy documents into every prompt, and conversations now take more turns to resolve because answers got vaguer. Total spend alone would have looked like a small rise. Cost per resolution made the problem, and its cause, obvious by Monday morning.

### Common mistakes

Common mistakes in monitoring. Averages instead of percentiles. Alerting on every single bad answer until on-call mutes the channel. Running uncalibrated judges on production traffic and trusting the numbers. And hard-coding token prices, so cost dashboards quietly go wrong when prices change. Quick question: if quality on one country's traffic dropped by ten points tomorrow, which dashboard would show it, and who would be alerted?

### Deeper: two-tier quality alerts

One level deeper on alert design. The team set two rules for quality. A warning when the judged pass rate falls below baseline minus three points for two consecutive hours. A page only when it falls below minus eight points, or any blocking category fires. Two thresholds, with a time window, cut noisy alerts dramatically while still catching real drops the same day.

### Watch me do it: the hourly monitor (illustrative)

Watch me do it with the hourly monitor job. It reads the last hour of exported traces: say four thousand two hundred requests. Step one, PII: the email check runs on every output and finds one hit. I open it: the bot repeated a customer's own email back when confirming a change. That is allowed by our policy when the customer is verified, so I refine the check to exclude the verified customer's own address, but the page still fires, which is right, because a human decides. Step two, latency: the job sorts latencies and takes the ninety-fifth percentile, about three point one seconds, within our objective. Step three, cost: it sums cost and divides by resolved conversations, not requests, and the result is in line with last week. Step four, quality: it samples five percent of traces that have a policy attached, about two hundred, and runs the calibrated policy judge. The pass rate is eighty-eight percent against a baseline of ninety-one, within our three-point tolerance, so it warns but does not page. Step five, drift check: I slice the judged sample by topic and see delivery questions at seventy-nine percent. That slice needs a look. The job prints one JSON report, the alerting system reads it, and failures reviewed today go into next month's dataset.

### Recap

Recap. Define SLOs per route using percentiles, track cost per successful outcome, run code checks on everything and calibrated judges on samples, and watch for input, knowledge and provider drift. Avoid alerting on single answers and never hard-code prices. Your next step: write three SLOs for your main LLM route, and schedule the monitoring job from the lesson to run hourly.

## Key takeaways

- Define SLIs and SLOs per route using percentiles: TTFT, latency, errors, cost per outcome, online quality, safety.
- Run code checks on all traffic, calibrated judges on samples, and use implicit feedback signals.
- Watch input, knowledge and provider drift; pin model versions and track refusals, finish reasons and length.
- Alert on sustained trends and blocking categories, and feed production failures into offline datasets.

## Try it

Write three SLOs for your main LLM route and schedule the monitoring job from this lesson hourly with PII paging and quality alerts.

- [Previous: Tracing LLM apps with OpenTelemetry GenAI conventions](https://optimizeall.com/learn/llm-evals-and-observability/tracing-with-opentelemetry-genai)
- [Next: A/B tests and online experiments for LLM features](https://optimizeall.com/learn/llm-evals-and-observability/ab-tests-and-online-experiments)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
