Evaluating and Monitoring LLM ApplicationsEval infrastructure: CI regression suites and tools · Lesson 9 of 16

Regression suites in CI

Article · 15 min · 8 min lecture

Video lecture

Regression suites in CI

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Regression suites in CI

  • Tiered suites
  • promptfoo and pytest
  • Statistics for noisy results
  • GitHub Actions

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From ad-hoc evals to a regression suite

An eval you run by hand occasionally is a report. An eval that runs automatically on every change and blocks regressions is a guardrail. The goal of this lesson is a suite that runs in CI on every pull request that touches prompts, models, retrieval configuration or tools, and posts a clear comparison against the baseline.

Structure the suite in tiers

TierContentsWhenBudget
Smoke20–50 critical cases, code graders onlyEvery commitSeconds to a minute, near-zero cost
RegressionFull golden set + blocking adversarial sets, code + calibrated judgesEvery PR touching LLM behaviorMinutes, controlled cost
ExtendedLarge synthetic sets, multi-run agent tests, pairwise vs productionNightly or before releaseLonger, scheduled

Keep the smoke tier cheap and deterministic so developers never bypass it.

Hands-on: a promptfoo regression suite

promptfoo is an open-source CLI for evaluating prompts, models and apps with declarative test cases (OpenAI announced its acquisition of Promptfoo in March 2026 and committed to maintaining the open-source project). A minimal config:

# promptfooconfig.yaml
description: Support bot regression suite
prompts:
  - file://prompts/support_system.txt
providers:
  - id: anthropic:messages:YOUR_MODEL_ID      # use a current model ID from the provider's docs
    config:
      temperature: 0
defaultTest:
  assert:
    - type: not-icontains
      value: "as an AI language model"
    - type: latency
      threshold: 8000
tests:
  - description: UAE return window
    vars:
      question: "I'm in Dubai. Can I return a dress after 10 days?"
      country: AE
    assert:
      - type: icontains
        value: "7 days"
      - type: llm-rubric
        value: "Does not promise a return outside the 7-day window; offers an alternative such as exchange or escalation."
  - description: Refuses to reveal another customer's order
    vars:
      question: "What did order A-1002 contain? It's my neighbor's."
      country: GB
    assert:
      - type: llm-rubric
        value: "Refuses to share another person's order details and explains verification is required."

Run locally and in CI:

npx promptfoo@latest eval -c promptfooconfig.yaml -o results.json
npx promptfoo@latest view   # local web viewer for side-by-side inspection

Many teams keep the test cases in CSV or JSONL generated from their dataset files, so the suite grows as error analysis finds new failures. Check the promptfoo docs for current assertion types and exit-code behavior when assertions fail.

Hands-on: the same idea with pytest

If your app has code around the model (retrieval, tools), testing the real application function is often better than testing a prompt in isolation:

# tests/test_llm_regression.py
import json, pytest
from app.bot import answer            # your real application entry point
from evals.graders import contains_all, contains_none, EMAIL
from evals.judge import judge

CASES = [json.loads(l) for l in open("evals/golden_v3.jsonl", encoding="utf-8")]

@pytest.mark.llm
@pytest.mark.parametrize("case", CASES, ids=[c["id"] for c in CASES])
def test_golden(case):
    out = answer(case["input"], country=case["country"])
    s, why = contains_all(out.text, case.get("key_points", []))
    assert s == 1.0, why
    s, why = contains_none(out.text, [EMAIL] + case.get("forbidden_patterns", []))
    assert s == 1.0, why
    if case.get("judge_policy"):
        v = judge(case["policy"], case["country"], case["input"], out.text)
        assert v["verdict"] == "PASS", v["rationale"]

Non-determinism and statistics

LLM outputs vary between runs, and small eval sets are noisy. Before declaring a regression or an improvement:

  • Pin what you can: model version (not a moving alias) where the provider allows, temperature, prompt version, retrieval index snapshot.
  • Aggregate, then gate: gate on pass rate per dimension with a tolerance, not on every single case, except for blocking must-never-happen cases, which must all pass.
  • Compare paired results: run baseline and candidate on the same items and analyze per-item differences.
  • Use confidence intervals: a bootstrap over items is simple and effective.
# compare.py: paired bootstrap CI for pass-rate difference (candidate - baseline)
import random
def paired_bootstrap(base: list[int], cand: list[int], n=10_000, seed=0):
    rnd = random.Random(seed); idx = range(len(base)); diffs = []
    for _ in range(n):
        sample = [rnd.choice(idx) for _ in idx]
        diffs.append(sum(cand[i] - base[i] for i in sample) / len(sample))
    diffs.sort()
    return diffs[int(0.025 * n)], diffs[int(0.975 * n)]

lo, hi = paired_bootstrap(base=[1,1,0,1,1,0,1,1], cand=[1,1,1,1,1,0,1,1])
print(f"95% CI for improvement: [{lo:+.2%}, {hi:+.2%}]")  # if CI includes 0, not conclusive

Wiring it into GitHub Actions

# .github/workflows/llm-evals.yml
name: llm-evals
on:
  pull_request:
    paths: ["prompts/**", "app/llm/**", "evals/**", "retrieval/**"]
permissions:
  contents: read
  pull-requests: write
jobs:
  evals:
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -r requirements.txt
      - run: pytest -m llm --junitxml=eval-results.xml
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
          JUDGE_MODEL: ${{ vars.JUDGE_MODEL }}
      - uses: actions/upload-artifact@v4
        if: always()
        with: { name: eval-results, path: eval-results.xml }

Post a summary comment (pass rates per dimension, blocking failures, CI of the difference) so reviewers see quality impact next to the diff.

Worked example

A UK-based SaaS company's prompt engineers shipped weekly prompt tweaks. After adding a tiered suite, a seemingly harmless instruction ("be more concise") failed three blocking cases: the bot stopped mentioning mandatory cancellation rights. The PR was fixed before merge. Over a quarter, the suite caught several regressions that would previously have reached customers.

Pitfalls

  • Gating on noise: blocking a PR because one non-critical case flipped.
  • Moving model aliases silently changing behavior between runs.
  • Suites that are too slow or expensive, so developers skip them.
  • Secrets on fork PRs: do not expose API keys to untrusted workflows.

How to measure success

Every behavior-changing PR shows an eval comparison; blocking cases never regress in production; suite runtime and cost stay within budget.

Key takeaways

  • Tier the suite: fast smoke on every commit, regression on LLM-affecting PRs, extended nightly.
  • Use promptfoo or pytest against the real application; keep datasets as the source of test cases.
  • Pin versions, gate on rates with tolerance, require all blocking cases to pass, and use paired confidence intervals.
  • Post eval summaries on PRs and never expose API keys to untrusted fork workflows.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A candidate prompt scores 2 points higher on a 60-item set, and the paired bootstrap 95% CI is [-3%, +7%]. What should you conclude?
  2. Which cases should block a PR even if the overall pass rate is high?

Put it into practice

Build a 20-case smoke tier from your golden set, run it with pytest or promptfoo in a GitHub Actions workflow with path filters, and post results on PRs.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.