---
title: "Evaluating and Monitoring LLM Applications — free course"
description: "Evals, LLM-as-judge, RAG and agent evaluation, CI regression suites and OpenTelemetry-based observability in production"
url: https://optimizeall.com/learn/llm-evals-and-observability
updated: 2026-10-05
---

AI · Advanced · 413 minutes · free · updated Sep 2026

# Evaluating and Monitoring LLM Applications

Evals, LLM-as-judge, RAG and agent evaluation, CI regression suites and OpenTelemetry-based observability in production

- **Lessons:** 16 in 6 modules
- **Video lectures:** 16 lectures, 133 minutes
- **Updated:** Sep 2026

## Tools you'll use

- promptfoo
- Inspect
- Langfuse
- LangSmith
- Braintrust
- Arize Phoenix
- OpenTelemetry
- Ragas
- DeepEval
- MLflow
- pytest
- GitHub Actions
- Anthropic API
- OpenAI API

[Start the course](https://optimizeall.com/learn/llm-evals-and-observability/why-evals-and-defining-success)

## About this course

Prompt tweaks, model upgrades and new documents can silently break an LLM product. This advanced course teaches the discipline that prevents it: evaluation and observability. You will define success with stakeholders, run error analysis to find how your system really fails, and build golden, production and synthetic datasets. You will design code-based metrics and calibrated LLM-as-judge graders, measure groundedness, safety and tone, and evaluate RAG pipelines and multi-step agents. You will turn evals into CI regression suites with statistical care, compare today's tools (promptfoo, Inspect, Langfuse, LangSmith, Braintrust, Arize Phoenix and more), trace production traffic with OpenTelemetry GenAI conventions, set cost and latency SLOs, run A/B tests and build human review workflows. The capstone is a working Python eval harness for a customer support bot, wired into CI and monitoring.

## What you will learn

- Define measurable success criteria and blocking failures for an LLM feature with stakeholders
- Run error analysis to build a failure taxonomy and targeted eval datasets
- Design code-based metrics and LLM-as-judge graders calibrated against human labels
- Evaluate RAG retrieval and generation, and agent trajectories and outcomes
- Run eval regression suites in CI with thresholds and statistical care
- Instrument LLM apps with OpenTelemetry GenAI conventions and set cost and latency SLOs
- Operate online evals, A/B tests and human review loops that feed back into datasets

## Before you start

- [Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
- [Latest AI Techniques: RAG, Tool Use, Agents & MCP](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp)

## Course content

### Foundations: success criteria, error analysis and datasets

Define what good looks like, find how your system actually fails through error analysis, and build golden, production and synthetic eval datasets.

- [Why evals, and defining success](https://optimizeall.com/learn/llm-evals-and-observability/why-evals-and-defining-success): 13 min
- [Error analysis and failure taxonomies](https://optimizeall.com/learn/llm-evals-and-observability/error-analysis-and-failure-taxonomies): 13 min
- [Building eval datasets: golden, production and synthetic](https://optimizeall.com/learn/llm-evals-and-observability/building-eval-datasets): 14 min

### Metrics and graders: code, judges and quality dimensions

Build deterministic graders first, design and calibrate LLM-as-judge graders, and measure groundedness, safety, refusals, privacy, fairness and tone.

- [Code-based metrics and graders](https://optimizeall.com/learn/llm-evals-and-observability/code-based-metrics): 13 min
- [LLM-as-judge: design, bias and calibration](https://optimizeall.com/learn/llm-evals-and-observability/llm-as-judge): 15 min
- [Groundedness, safety, privacy and fairness metrics](https://optimizeall.com/learn/llm-evals-and-observability/groundedness-safety-and-quality): 14 min

### Evaluating RAG pipelines and agents

Evaluate retrieval and generation separately with a 2×2 diagnosis, and evaluate agents on outcomes, trajectories and efficiency in sandboxes with repeated runs.

- [Evaluating RAG: retrieval, generation and abstention](https://optimizeall.com/learn/llm-evals-and-observability/evaluating-rag): 15 min
- [Evaluating agents: outcomes, trajectories and efficiency](https://optimizeall.com/learn/llm-evals-and-observability/evaluating-agents-and-trajectories): 15 min

### Eval infrastructure: CI regression suites and tools

Turn evals into tiered regression suites in CI with statistical care, and choose tools from the 2026 landscape while keeping datasets and instrumentation portable.

- [Regression suites in CI](https://optimizeall.com/learn/llm-evals-and-observability/regression-suites-in-ci): 15 min
- [The 2026 evals and observability tools landscape](https://optimizeall.com/learn/llm-evals-and-observability/evals-tools-landscape): 13 min

### Production observability, monitoring and experimentation

Trace LLM apps with OpenTelemetry GenAI conventions, run online monitoring with SLOs and drift detection, run safe A/B tests, and operate human review loops.

- [Tracing LLM apps with OpenTelemetry GenAI conventions](https://optimizeall.com/learn/llm-evals-and-observability/tracing-with-opentelemetry-genai): 15 min
- [Online monitoring, SLOs and drift](https://optimizeall.com/learn/llm-evals-and-observability/online-monitoring-and-slos): 14 min
- [A/B tests and online experiments for LLM features](https://optimizeall.com/learn/llm-evals-and-observability/ab-tests-and-online-experiments): 13 min
- [Human review workflows and feedback loops](https://optimizeall.com/learn/llm-evals-and-observability/human-review-workflows): 13 min

### Capstone: an eval harness for a support bot

Build a Python eval harness for a multi-country support bot, compare a cheaper candidate rigorously, wire it into CI and monitoring, and write the ship decision.

- [Capstone part 1: dataset, interface and harness](https://optimizeall.com/learn/llm-evals-and-observability/capstone-build-the-harness): 20 min
- [Capstone part 2: comparison, CI, monitoring and the ship decision](https://optimizeall.com/learn/llm-evals-and-observability/capstone-ci-monitoring-and-decision): 20 min

## Certificate: Certified LLM Evaluation Engineer

The holder can measure and monitor LLM applications rigorously: defining success criteria, running error analysis, building datasets, designing code-based and calibrated LLM-as-judge metrics, evaluating RAG and agents, running regression suites in CI, tracing production with OpenTelemetry, setting cost and latency SLOs, and operating A/B tests and human review loops.

- **Final assessment:** 30 questions, 45 minutes
- **Passing score:** 80%
