AI · Advanced · 413 minutes · free · updated Sep 2026
Evaluating and Monitoring LLM Applications
Evals, LLM-as-judge, RAG and agent evaluation, CI regression suites and OpenTelemetry-based observability in production
- Lessons
- 16 in 6 modules
- Video lectures
- 16 lectures, 133 minutes
- Updated
- Sep 2026
Tools you'll use
- promptfoo
- Inspect
- Langfuse
- LangSmith
- Braintrust
- Arize Phoenix
- OpenTelemetry
- Ragas
- DeepEval
- MLflow
- pytest
- GitHub Actions
- Anthropic API
- OpenAI API
About this course
Prompt tweaks, model upgrades and new documents can silently break an LLM product. This advanced course teaches the discipline that prevents it: evaluation and observability. You will define success with stakeholders, run error analysis to find how your system really fails, and build golden, production and synthetic datasets. You will design code-based metrics and calibrated LLM-as-judge graders, measure groundedness, safety and tone, and evaluate RAG pipelines and multi-step agents. You will turn evals into CI regression suites with statistical care, compare today's tools (promptfoo, Inspect, Langfuse, LangSmith, Braintrust, Arize Phoenix and more), trace production traffic with OpenTelemetry GenAI conventions, set cost and latency SLOs, run A/B tests and build human review workflows. The capstone is a working Python eval harness for a customer support bot, wired into CI and monitoring.
What you will learn
- Define measurable success criteria and blocking failures for an LLM feature with stakeholders
- Run error analysis to build a failure taxonomy and targeted eval datasets
- Design code-based metrics and LLM-as-judge graders calibrated against human labels
- Evaluate RAG retrieval and generation, and agent trajectories and outcomes
- Run eval regression suites in CI with thresholds and statistical care
- Instrument LLM apps with OpenTelemetry GenAI conventions and set cost and latency SLOs
- Operate online evals, A/B tests and human review loops that feed back into datasets
Before you start
- Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API
- Latest AI Techniques: RAG, Tool Use, Agents & MCP
Course content
Foundations: success criteria, error analysis and datasets
Define what good looks like, find how your system actually fails through error analysis, and build golden, production and synthetic eval datasets.
- Why evals, and defining success — 13 min
- Error analysis and failure taxonomies — 13 min
- Building eval datasets: golden, production and synthetic — 14 min
Metrics and graders: code, judges and quality dimensions
Build deterministic graders first, design and calibrate LLM-as-judge graders, and measure groundedness, safety, refusals, privacy, fairness and tone.
- Code-based metrics and graders — 13 min
- LLM-as-judge: design, bias and calibration — 15 min
- Groundedness, safety, privacy and fairness metrics — 14 min
Evaluating RAG pipelines and agents
Evaluate retrieval and generation separately with a 2×2 diagnosis, and evaluate agents on outcomes, trajectories and efficiency in sandboxes with repeated runs.
- Evaluating RAG: retrieval, generation and abstention — 15 min
- Evaluating agents: outcomes, trajectories and efficiency — 15 min
Eval infrastructure: CI regression suites and tools
Turn evals into tiered regression suites in CI with statistical care, and choose tools from the 2026 landscape while keeping datasets and instrumentation portable.
- Regression suites in CI — 15 min
- The 2026 evals and observability tools landscape — 13 min
Production observability, monitoring and experimentation
Trace LLM apps with OpenTelemetry GenAI conventions, run online monitoring with SLOs and drift detection, run safe A/B tests, and operate human review loops.
- Tracing LLM apps with OpenTelemetry GenAI conventions — 15 min
- Online monitoring, SLOs and drift — 14 min
- A/B tests and online experiments for LLM features — 13 min
- Human review workflows and feedback loops — 13 min
Capstone: an eval harness for a support bot
Build a Python eval harness for a multi-country support bot, compare a cheaper candidate rigorously, wire it into CI and monitoring, and write the ship decision.
- Capstone part 1: dataset, interface and harness — 20 min
- Capstone part 2: comparison, CI, monitoring and the ship decision — 20 min
Certificate: Certified LLM Evaluation Engineer
The holder can measure and monitor LLM applications rigorously: defining success criteria, running error analysis, building datasets, designing code-based and calibrated LLM-as-judge metrics, evaluating RAG and agents, running regression suites in CI, tracing production with OpenTelemetry, setting cost and latency SLOs, and operating A/B tests and human review loops.
- Final assessment
- 30 questions, 45 minutes
- Passing score
- 80%