---
title: "Capstone part 2: comparison, CI, monitoring and the ship…"
description: "Part 2: from harness to decision In Part 1 you built the dataset, the bot interface, the harness and a baseline. Now you will evaluate the candidate…"
url: https://optimizeall.com/learn/llm-evals-and-observability/capstone-ci-monitoring-and-decision
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Capstone: an eval harness for a support bot · lesson 16 of 16 · 20 min

# Capstone part 2: comparison, CI, monitoring and the ship decision

## Part 2: from harness to decision

In Part 1 you built the dataset, the bot interface, the harness and a baseline. Now you will evaluate the candidate (cheaper model plus new prompt), wire the harness into CI, add production monitoring, and write the decision report leadership asked for.

## Step 5: evaluate the candidate

```bash
python -m evals.harness --model "$BASELINE_MODEL"  --prompt v12 --out results_base.json
python -m evals.harness --model "$CANDIDATE_MODEL" --prompt v13 --out results_cand.json
```

Run each configuration three times (or more for small datasets). Then compare paired results per case:

```python
# evals/compare.py: paired comparison with bootstrap CI and slice deltas
import json, random, sys

def load(p):
    d = json.load(open(p)); return {r["id"]: r for r in d["rows"]}, d["summary"]

base, bs = load(sys.argv[1]); cand, cs = load(sys.argv[2])
ids = sorted(set(base) & set(cand))
b = [int(base[i]["passed"]) for i in ids]; c = [int(cand[i]["passed"]) for i in ids]

rnd = random.Random(0); diffs = []
for _ in range(5000):
    s = [rnd.randrange(len(ids)) for _ in ids]
    diffs.append(sum(c[k] - b[k] for k in s) / len(s))
diffs.sort()
print(f"pass-rate delta {sum(c)/len(c) - sum(b)/len(b):+.1%}, 95% CI [{diffs[125]:+.1%}, {diffs[4874]:+.1%}]")
print("newly failing:", [i for i in ids if base[i]["passed"] and not cand[i]["passed"]])
print("newly passing:", [i for i in ids if cand[i]["passed"] and not base[i]["passed"]])
for k in sorted(set(bs["by_slice"]) | set(cs["by_slice"])):
    print(f"{k:18} base {bs['by_slice'].get(k, 0):.0%}  cand {cs['by_slice'].get(k, 0):.0%}")
print("cand blocking failures:", cs["blocking_failures"])
print(f"p95 latency ms: base {bs['p95_latency_ms']:.0f} cand {cs['p95_latency_ms']:.0f}")
print(f"mean tokens: base {bs['mean_tokens']:.0f} cand {cs['mean_tokens']:.0f}")
```

**Read the newly failing cases yourself.** Aggregate numbers tell you whether to look; individual cases tell you what changed.

## Step 6: wire into CI

```yaml
# .github/workflows/evals.yml
name: evals
on:
  pull_request:
    paths: ["app/**", "evals/**", "prompts/**"]
permissions:
  contents: read
jobs:
  eval:
    if: github.event.pull_request.head.repo.full_name == github.repository
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with: { python-version: "3.12" }
      - run: pip install -r requirements.txt
      - run: python -m evals.harness --model "${{ vars.APP_MODEL }}" --prompt "${{ vars.PROMPT_VERSION }}" --out results.json
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
          JUDGE_MODEL: ${{ vars.JUDGE_MODEL }}
      - run: python -m evals.compare evals/baseline.json results.json | tee compare.txt
      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: eval-report
          path: |
            results.json
            compare.txt
```

The harness exits non-zero on blocking failures, so the check fails the PR automatically. Add a tolerance rule in `compare.py` (for example fail if any slice drops by more than your agreed threshold with a CI excluding zero).

## Step 7: production monitoring

Instrument `answer()` with OpenTelemetry GenAI spans (Module 5), adding `souq.prompt_version`, `souq.country` and `souq.topic` attributes. Then schedule the hourly monitor job:

- PII and forbidden-pattern checks on 100% of responses; page on any hit.
- `lookup_order` called without verification: page on any occurrence (blocking trajectory rule).
- Policy judge on a 5% sample stratified by country; alert on a sustained drop below baseline minus tolerance.
- p95 latency and cost per resolved conversation per country.
- Weekly: 50 items into the human review queue (random + flagged + judge failures), double-label 10.

## Step 8: the decision report

```markdown
# Decision: switch support bot to <candidate model> + prompt v13
## Recommendation: SHIP / SHIP WITH CONDITIONS / DO NOT SHIP
## Evidence
- Overall pass rate: base __% vs cand __% (delta __, 95% CI [__, __]) over 3 runs each
- Blocking failures: base __ / cand __ (list)
- Slices with material change: __ (e.g. AE/returns −__ points; SA/delivery +__)
- Newly failing cases reviewed: __ (summary of root causes)
- Latency p95: __ → __ ms; mean tokens __ → __; estimated cost per resolved conversation __ → __
## Risks and mitigations
## Rollout plan
- Shadow 1 week → canary 5% with auto-rollback on guardrails → A/B by account (primary: resolution; guardrails: escalations, complaints, PII, latency)
## Follow-ups
- New eval cases added from this analysis: __
```

## Assessment rubric

| Criterion | Excellent |
|---|---|
| Dataset | ≥ 60 cases, all slices, sources documented, blocking set present |
| Harness | Grades text, tools and abstention; robust to errors; sliced summaries; non-zero exit on blocking |
| Statistics | Multiple runs; paired comparison with CI; newly failing cases read and explained |
| CI | Path-filtered workflow, least privilege, fork-safe, artifacts saved |
| Monitoring | GenAI spans with custom attributes; 100% code checks; sampled judge; review queue |
| Decision | Clear recommendation tied to evidence, with a staged rollout plan |

## Pitfalls

- **Deciding from the overall pass rate alone.**
- **Skipping the newly failing cases.**
- **Shipping without a rollback plan.**

## What next

Keep the loop alive: every production failure found by monitoring or review becomes a new case in `dataset_v2.jsonl`, and the baseline is refreshed at each release.

## Video lecture: Capstone part 2: comparison, CI, monitoring and the ship decision

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Capstone part 2: from harness to decision
2. Analogy: from trial to treatment
3. Compare rigorously
4. Read the newly failing cases
5. CI
6. Monitoring
7. Decision report
8. Rubric
9. Pitfalls + keep the loop alive
10. Example decision (illustrative)
11. Common mistakes
12. Deeper: a slice tolerance rule
13. Watch me do it: compare → decision (illustrative)
14. Recap

## Lecture transcript

### Capstone part 2: from harness to decision

In part one you built the dataset, the harness and a baseline. Now comes the moment of truth. Is the cheaper model with the new prompt safe to ship? In this lecture you will run the candidate, compare it rigorously with the baseline, wire the harness into CI, set up production monitoring, and write the decision report that leadership actually needs.

### Analogy: from trial to treatment

An analogy for this part. A good medical trial does not end with the lab result. It ends with a decision, a monitoring plan and a way to stop if side effects appear. Part two is the same. Your comparison is the trial result, CI is the protocol that every future change must follow, monitoring is the post-launch surveillance, and the decision report is where you recommend a treatment, with its conditions and its stopping rules.

### Compare rigorously

Run both configurations through the harness, three times each. Then compare them case by case. The comparison script in the lesson computes the change in pass rate with a paired bootstrap confidence interval, lists the cases that newly fail and newly pass, shows pass rates for every country and topic slice, reports blocking failures, and compares p95 latency and token usage.

### Read the newly failing cases

Here is the most important instruction in the whole capstone. Read the newly failing cases yourself. Aggregate numbers tell you whether to look. Individual cases tell you what actually changed. Maybe the new prompt drops the seven-day limit for Emirati returns. Maybe the cheaper model handles code-switched Urdu worse. Those are the details that turn a number into a decision, and into better tests.

### CI

Next, wire it into CI. The workflow triggers on pull requests that touch the app, evals or prompts, skips forks, runs with read-only permissions, reads keys from secrets and the model names from repository variables, runs the harness and the comparison, and uploads the report as an artifact. Because the harness exits non-zero on any blocking failure, the pull request fails automatically. Add a tolerance rule for slice regressions too.

### Monitoring

Then production monitoring. Instrument the answer function with OpenTelemetry GenAI spans, plus attributes for prompt version, country and topic. Run the hourly monitor: PII and forbidden patterns on every response, paging on any hit. The trajectory rule that lookup order is never called without verification, also paging. The policy judge on a five percent sample stratified by country. Latency and cost per resolved conversation by country. And fifty items a week into human review, with ten double-labeled.

### Decision report

Finally, the decision report. Lead with a clear recommendation: ship, ship with conditions, or do not ship. Then the evidence: overall pass rates with the confidence interval over three runs each, blocking failures, slices with material changes, the root causes of newly failing cases, and latency, token and cost changes. Then risks and mitigations, and a staged rollout plan: shadow for a week, a five percent canary with automatic rollback, then an A/B test by account.

### Rubric

You will be assessed on six things. The dataset: at least sixty cases, all slices, documented sources and a blocking set. The harness: grades text, tools and abstention, handles errors and exits correctly. Statistics: multiple runs, paired confidence intervals, and newly failing cases explained. CI: filtered, least privilege and fork safe. Monitoring: spans, full-coverage code checks, a sampled judge and a review queue. And the decision: a clear recommendation tied to evidence with a staged rollout.

### Pitfalls + keep the loop alive

Three pitfalls to avoid. Deciding from the overall pass rate alone, when a slice or a blocking case tells a different story. Skipping the newly failing cases. And shipping without a rollback plan. And remember that the loop never ends. Every failure found in production by monitoring or review becomes a new case in dataset version two, and you refresh the baseline at every release.

### Example decision (illustrative)

A simple example of turning numbers into a decision, with illustrative results. Overall: candidate eighty-seven percent versus baseline eighty-six, confidence interval from minus three to plus five. Blocking failures: none. Slices: UAE returns dropped from ninety to seventy-eight, and the newly failing cases all involve the seven-day window being stated as fourteen. Cost per conversation: down by roughly forty percent. Recommendation: ship with conditions. Fix the UAE return wording and add three cases, then shadow for a week, then a five percent canary with the UAE returns slice as a guardrail.

### Common mistakes

Common mistakes at this stage. Presenting only the overall number to leadership. Skipping the newly failing cases because the average looked fine. Shipping to everyone at once with no rollback plan. And forgetting to add today's discoveries to dataset version two, so the same failure can sneak back next quarter. Try this now: draft the one-line recommendation for your own capstone before you write anything else, then make sure every sentence in the report supports or qualifies it.

### Deeper: a slice tolerance rule

One level deeper on the tolerance rule in compare. Fail the pull request if any slice with at least ten cases drops by more than eight points and the bootstrap interval for that slice excludes zero. Smaller slices only warn. That rule would have blocked the AE returns regression automatically, while ignoring noise in tiny slices.

### Watch me do it: compare → decision (illustrative)

Watch me do it with the comparison script. I run the harness three times for the baseline and three times for the candidate, then run compare on the pair of results files. First output line: the pass-rate delta, plus one point, with a ninety-five percent interval from minus three to plus five. Inconclusive overall. Second: newly failing cases, four of them, all IDs starting with ret AE. Third: newly passing, three cases, mostly code-switched Urdu questions, interesting. Fourth: the slice table. AE returns drops from ninety to seventy-eight percent. PK delivery rises. Everything else is within a few points. Fifth: blocking failures for the candidate, none. Sixth: p95 latency slightly lower, and mean tokens much lower. Now I read the four newly failing cases. In all four, the candidate says fourteen days for a UAE return. I check the new prompt: it contains an example answer taken from the UK policy, and the cheaper model copies it too literally. So the recommendation writes itself: ship with conditions. Replace the UK example with a country-neutral one, add three more UAE returns cases, rerun the harness, then shadow for a week and canary at five percent with AE returns as a guardrail. I paste those facts into the decision report template, and the whole report fits on one page.

### Recap

Recap. Evaluate the candidate with repeated runs and paired statistics, read every newly failing case, gate pull requests in CI, monitor production with code checks, sampled judges and human review, and make a clear, evidence-based recommendation with a staged rollout. Your next step: complete the capstone, and present your decision report to a colleague as if they were the leadership team.

## Key takeaways

- Compare candidate and baseline with repeated runs, paired bootstrap CIs, slice deltas and blocking failures.
- Read every newly failing case; aggregate numbers only say where to look.
- Gate PRs with a fork-safe, least-privilege CI workflow and monitor production with spans, code checks, sampled judges and review.
- Write a recommendation tied to evidence with a staged rollout and rollback plan.

## Try it

Run the candidate comparison, wire the CI workflow, set up the monitor job, and write the decision report using the template.

- [Previous: Capstone part 1: dataset, interface and harness](https://optimizeall.com/learn/llm-evals-and-observability/capstone-build-the-harness)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
