Evaluating and Monitoring LLM ApplicationsCapstone: an eval harness for a support bot · Lesson 16 of 16

Capstone part 2: comparison, CI, monitoring and the ship decision

Article · 20 min · 8 min lecture

Video lecture

Capstone part 2: comparison, CI, monitoring and the ship decision

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Capstone part 2: from harness to decision

  • Evaluate the candidate
  • CI
  • Monitoring
  • Decision report

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Part 2: from harness to decision

In Part 1 you built the dataset, the bot interface, the harness and a baseline. Now you will evaluate the candidate (cheaper model plus new prompt), wire the harness into CI, add production monitoring, and write the decision report leadership asked for.

Step 5: evaluate the candidate

python -m evals.harness --model "$BASELINE_MODEL"  --prompt v12 --out results_base.json
python -m evals.harness --model "$CANDIDATE_MODEL" --prompt v13 --out results_cand.json

Run each configuration three times (or more for small datasets). Then compare paired results per case:

# evals/compare.py: paired comparison with bootstrap CI and slice deltas
import json, random, sys

def load(p):
    d = json.load(open(p)); return {r["id"]: r for r in d["rows"]}, d["summary"]

base, bs = load(sys.argv[1]); cand, cs = load(sys.argv[2])
ids = sorted(set(base) & set(cand))
b = [int(base[i]["passed"]) for i in ids]; c = [int(cand[i]["passed"]) for i in ids]

rnd = random.Random(0); diffs = []
for _ in range(5000):
    s = [rnd.randrange(len(ids)) for _ in ids]
    diffs.append(sum(c[k] - b[k] for k in s) / len(s))
diffs.sort()
print(f"pass-rate delta {sum(c)/len(c) - sum(b)/len(b):+.1%}, 95% CI [{diffs[125]:+.1%}, {diffs[4874]:+.1%}]")
print("newly failing:", [i for i in ids if base[i]["passed"] and not cand[i]["passed"]])
print("newly passing:", [i for i in ids if cand[i]["passed"] and not base[i]["passed"]])
for k in sorted(set(bs["by_slice"]) | set(cs["by_slice"])):
    print(f"{k:18} base {bs['by_slice'].get(k, 0):.0%}  cand {cs['by_slice'].get(k, 0):.0%}")
print("cand blocking failures:", cs["blocking_failures"])
print(f"p95 latency ms: base {bs['p95_latency_ms']:.0f} cand {cs['p95_latency_ms']:.0f}")
print(f"mean tokens: base {bs['mean_tokens']:.0f} cand {cs['mean_tokens']:.0f}")

Read the newly failing cases yourself. Aggregate numbers tell you whether to look; individual cases tell you what changed.

Step 6: wire into CI

# .github/workflows/evals.yml
name: evals
on:
  pull_request:
    paths: ["app/**", "evals/**", "prompts/**"]
permissions:
  contents: read
jobs:
  eval:
    if: github.event.pull_request.head.repo.full_name == github.repository
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -r requirements.txt
      - run: python -m evals.harness --model "${{ vars.APP_MODEL }}" --prompt "${{ vars.PROMPT_VERSION }}" --out results.json
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
          JUDGE_MODEL: ${{ vars.JUDGE_MODEL }}
      - run: python -m evals.compare evals/baseline.json results.json | tee compare.txt
      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: eval-report
          path: |
            results.json
            compare.txt

The harness exits non-zero on blocking failures, so the check fails the PR automatically. Add a tolerance rule in compare.py (for example fail if any slice drops by more than your agreed threshold with a CI excluding zero).

Step 7: production monitoring

Instrument answer() with OpenTelemetry GenAI spans (Module 5), adding souq.prompt_version, souq.country and souq.topic attributes. Then schedule the hourly monitor job:

  • PII and forbidden-pattern checks on 100% of responses; page on any hit.
  • lookup_order called without verification: page on any occurrence (blocking trajectory rule).
  • Policy judge on a 5% sample stratified by country; alert on a sustained drop below baseline minus tolerance.
  • p95 latency and cost per resolved conversation per country.
  • Weekly: 50 items into the human review queue (random + flagged + judge failures), double-label 10.

Step 8: the decision report

# Decision: switch support bot to <candidate model> + prompt v13
## Recommendation: SHIP / SHIP WITH CONDITIONS / DO NOT SHIP
## Evidence
- Overall pass rate: base __% vs cand __% (delta __, 95% CI [__, __]) over 3 runs each
- Blocking failures: base __ / cand __ (list)
- Slices with material change: __ (e.g. AE/returns −__ points; SA/delivery +__)
- Newly failing cases reviewed: __ (summary of root causes)
- Latency p95: __ → __ ms; mean tokens __ → __; estimated cost per resolved conversation __ → __
## Risks and mitigations
## Rollout plan
- Shadow 1 week → canary 5% with auto-rollback on guardrails → A/B by account (primary: resolution; guardrails: escalations, complaints, PII, latency)
## Follow-ups
- New eval cases added from this analysis: __

Assessment rubric

CriterionExcellent
Dataset≥ 60 cases, all slices, sources documented, blocking set present
HarnessGrades text, tools and abstention; robust to errors; sliced summaries; non-zero exit on blocking
StatisticsMultiple runs; paired comparison with CI; newly failing cases read and explained
CIPath-filtered workflow, least privilege, fork-safe, artifacts saved
MonitoringGenAI spans with custom attributes; 100% code checks; sampled judge; review queue
DecisionClear recommendation tied to evidence, with a staged rollout plan

Pitfalls

  • Deciding from the overall pass rate alone.
  • Skipping the newly failing cases.
  • Shipping without a rollback plan.

What next

Keep the loop alive: every production failure found by monitoring or review becomes a new case in dataset_v2.jsonl, and the baseline is refreshed at each release.

Key takeaways

  • Compare candidate and baseline with repeated runs, paired bootstrap CIs, slice deltas and blocking failures.
  • Read every newly failing case; aggregate numbers only say where to look.
  • Gate PRs with a fork-safe, least-privilege CI workflow and monitor production with spans, code checks, sampled judges and review.
  • Write a recommendation tied to evidence with a staged rollout and rollback plan.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. The candidate's overall pass rate is 1 point higher, but the AE/returns slice drops 12 points. What should the report do?
  2. Which rollout plan is most appropriate after offline evals pass?

Put it into practice

Run the candidate comparison, wire the CI workflow, set up the monitor job, and write the decision report using the template.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.