Evaluating and Monitoring LLM ApplicationsProduction observability, monitoring and experimentation · Lesson 12 of 16

Online monitoring, SLOs and drift

Article · 14 min · 9 min lecture

Video lecture

Online monitoring, SLOs and drift

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Online monitoring, SLOs and drift

  • SLIs and SLOs for LLM features
  • Cost and online quality
  • Detecting drift

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Offline evals are not enough

Offline evals test the cases you thought of. Production brings new users, new phrasing, new documents, provider changes and traffic spikes. Online monitoring watches quality, safety, latency and cost continuously on real traffic, and turns surprises into alerts and new test cases.

Service level objectives for LLM features

Borrow SRE practice: define SLIs (indicators) and SLOs (targets) that reflect user experience.

SLIExample SLO (illustrative; set your own)
Time to first token (streaming)p95 below an agreed threshold during business hours
End-to-end latency per requestp95 and p99 targets per route
Error rate (provider errors, timeouts, tool failures)Below an agreed percentage over 28 days
Cost per conversation or per resolved ticketBelow budget, with alert at 80% of budget
Online quality pass rate (sampled judges)At or above offline baseline minus a tolerance
Safety violationsZero on blocking categories; any occurrence pages on-call

Latency for LLMs depends on input and output length, so track tokens per request alongside latency, and consider separate SLOs for short and long tasks. Use percentiles, not averages.

Cost monitoring

  • Compute cost from token attributes and your current price sheet (keep prices in config, not code; they change).
  • Break down by route, customer tier, model and prompt version.
  • Alert on cost per successful outcome, not just total spend: a cheaper model that fails more may cost more per resolved ticket.
  • Watch for runaway loops in agents (step counts, repeated tool calls) and set hard limits.

Online quality evaluation

You cannot label all production traffic, but you can:

  1. Run code checks on 100% of traffic: format validity, forbidden patterns, PII detectors, tool-call rules, refusal markers, length limits.
  2. Run LLM judges on a sample (for example a few percent, or stratified by route and risk), using the same calibrated judges as offline.
  3. Use implicit signals: thumbs up/down, escalation to human, user rephrasing the same question, conversation abandonment, repeated contacts within 24 hours.
  4. Compare with the offline baseline; alert on sustained drops, not single bad answers.

Detecting drift

Things drift even when you change nothing:

  • Input drift: new topics, languages, seasonal questions (Ramadan and Eid shopping, back-to-school, Black Friday, Diwali or White Friday sales in the Gulf). Monitor topic clusters and language mix.
  • Knowledge drift: documents change; retrieval indexes go stale.
  • Provider drift: model aliases updated, behavior tweaks, new safety filters. Pin versions where possible and watch finish reasons, refusal rates and output length.

Hands-on: an online monitor job

# online_monitor.py: hourly job over recent traces exported as JSONL (adapt to your backend's API)
import json, random, statistics
from evals.graders import contains_none, EMAIL
from evals.judge import judge

SAMPLE_RATE = 0.05
traces = [json.loads(l) for l in open("last_hour_traces.jsonl", encoding="utf-8")]

pii_hits = [t["id"] for t in traces if contains_none(t["output"], [EMAIL])[0] == 0.0]
latencies = sorted(t["latency_ms"] for t in traces)
p95 = latencies[int(0.95 * (len(latencies) - 1))] if latencies else 0
cost = sum(t["cost_usd"] for t in traces)
resolved = sum(1 for t in traces if t.get("outcome") == "resolved") or 1

sample = [t for t in traces if random.random() < SAMPLE_RATE and t.get("policy")]
verdicts = [judge(t["policy"], t["country"], t["input"], t["output"])["verdict"] for t in sample]
pass_rate = verdicts.count("PASS") / max(1, len([v for v in verdicts if v != "ERROR"]))

report = {"n": len(traces), "pii_hits": pii_hits, "p95_ms": p95,
          "cost_per_resolved": cost / resolved, "judge_pass_rate": pass_rate, "judged": len(sample)}
print(json.dumps(report))
# send to your alerting system; page on any PII hit, warn on pass_rate below baseline - tolerance

Failures found this way should be added to the offline dataset after review.

Worked example

A UK retailer's assistant held steady for months, then online judge pass rate for delivery questions dropped over a week. Offline evals still passed. Trace slicing showed a surge of questions about a new "nominated day" delivery option that did not exist in the knowledge base yet. The assistant was guessing. The fix was content (add the policy), plus 20 new eval cases and an abstention check for unknown delivery options.

Pitfalls

  • Averages instead of percentiles.
  • Alerting on single bad answers (alert fatigue) instead of sustained trends and blocking categories.
  • Judging production with uncalibrated judges.
  • Hard-coded prices that silently go stale.

How to measure success

SLOs defined and visible per route, alerts that fire rarely and meaningfully, and a steady flow of production failures into offline datasets.

Key takeaways

  • Define SLIs and SLOs per route using percentiles: TTFT, latency, errors, cost per outcome, online quality, safety.
  • Run code checks on all traffic, calibrated judges on samples, and use implicit feedback signals.
  • Watch input, knowledge and provider drift; pin model versions and track refusals, finish reasons and length.
  • Alert on sustained trends and blocking categories, and feed production failures into offline datasets.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A cheaper model cuts cost per call by 30% but resolved-ticket rate drops and escalations rise. Which metric best captures the trade-off?
  2. Offline evals pass, but online judge pass rate drops for one topic. What is a likely cause?

Put it into practice

Write three SLOs for your main LLM route and schedule the monitoring job from this lesson hourly with PII paging and quality alerts.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.