Skip to content

AI · Advanced · 413 minutes · free · updated Sep 2026

Evaluating and Monitoring LLM Applications

Evals, LLM-as-judge, RAG and agent evaluation, CI regression suites and OpenTelemetry-based observability in production

Lessons
16 in 6 modules
Video lectures
16 lectures, 133 minutes
Updated
Sep 2026

Tools you'll use

  • promptfoo
  • Inspect
  • Langfuse
  • LangSmith
  • Braintrust
  • Arize Phoenix
  • OpenTelemetry
  • Ragas
  • DeepEval
  • MLflow
  • pytest
  • GitHub Actions
  • Anthropic API
  • OpenAI API

Start the course

About this course

Prompt tweaks, model upgrades and new documents can silently break an LLM product. This advanced course teaches the discipline that prevents it: evaluation and observability. You will define success with stakeholders, run error analysis to find how your system really fails, and build golden, production and synthetic datasets. You will design code-based metrics and calibrated LLM-as-judge graders, measure groundedness, safety and tone, and evaluate RAG pipelines and multi-step agents. You will turn evals into CI regression suites with statistical care, compare today's tools (promptfoo, Inspect, Langfuse, LangSmith, Braintrust, Arize Phoenix and more), trace production traffic with OpenTelemetry GenAI conventions, set cost and latency SLOs, run A/B tests and build human review workflows. The capstone is a working Python eval harness for a customer support bot, wired into CI and monitoring.

What you will learn

  • Define measurable success criteria and blocking failures for an LLM feature with stakeholders
  • Run error analysis to build a failure taxonomy and targeted eval datasets
  • Design code-based metrics and LLM-as-judge graders calibrated against human labels
  • Evaluate RAG retrieval and generation, and agent trajectories and outcomes
  • Run eval regression suites in CI with thresholds and statistical care
  • Instrument LLM apps with OpenTelemetry GenAI conventions and set cost and latency SLOs
  • Operate online evals, A/B tests and human review loops that feed back into datasets

Before you start

Course content

Foundations: success criteria, error analysis and datasets

Define what good looks like, find how your system actually fails through error analysis, and build golden, production and synthetic eval datasets.

Metrics and graders: code, judges and quality dimensions

Build deterministic graders first, design and calibrate LLM-as-judge graders, and measure groundedness, safety, refusals, privacy, fairness and tone.

Evaluating RAG pipelines and agents

Evaluate retrieval and generation separately with a 2×2 diagnosis, and evaluate agents on outcomes, trajectories and efficiency in sandboxes with repeated runs.

Eval infrastructure: CI regression suites and tools

Turn evals into tiered regression suites in CI with statistical care, and choose tools from the 2026 landscape while keeping datasets and instrumentation portable.

Production observability, monitoring and experimentation

Trace LLM apps with OpenTelemetry GenAI conventions, run online monitoring with SLOs and drift detection, run safe A/B tests, and operate human review loops.

Capstone: an eval harness for a support bot

Build a Python eval harness for a multi-country support bot, compare a cheaper candidate rigorously, wire it into CI and monitoring, and write the ship decision.

Certificate: Certified LLM Evaluation Engineer

The holder can measure and monitor LLM applications rigorously: defining success criteria, running error analysis, building datasets, designing code-based and calibrated LLM-as-judge metrics, evaluating RAG and agents, running regression suites in CI, tracing production with OpenTelemetry, setting cost and latency SLOs, and operating A/B tests and human review loops.

Final assessment
30 questions, 45 minutes
Passing score
80%