Evaluating and Monitoring LLM Applications · Production observability, monitoring and experimentation · lesson 13 of 16 · 13 min
A/B tests and online experiments for LLM features
When offline evals cannot decide
Offline evals tell you whether a change is safe and roughly better. They cannot fully tell you whether users will resolve more issues, convert more, or come back. For that you run online experiments: expose a random share of real traffic to the new variant and compare outcomes.
Stages of rollout
- Offline evals pass (blocking cases, no regressions beyond tolerance).
- Shadow mode: run the new variant on live inputs in parallel without showing its output to users; compare with offline graders and judges. Zero user risk, real input distribution. Mind cost and privacy.
- Canary: a small share of traffic (for example a few percent) sees the new variant; watch guardrail metrics closely; auto-rollback on breach.
- A/B test: a planned split (often 50/50) for a pre-computed duration to measure the primary metric.
- Full rollout, with monitoring continuing.
Designing an A/B test for an LLM feature
- Hypothesis: "Adding order-status tool calls will increase first-contact resolution without raising escalations."
- Randomization unit: usually the user (or account), not the request, so the same person sees a consistent experience and outcomes are independent.
- Primary metric: one metric that reflects value: resolution rate, conversion, task completion, retention.
- Guardrail metrics: must not get worse: safety violations, escalations, complaints, latency p95, cost per conversation, refund rate.
- Duration and sample size: computed before starting; cover full weekly cycles.
- Analysis plan: written in advance to avoid cherry-picking.
Sample size, roughly
A widely used rule of thumb for comparing two means (about 80% power at 5% significance) is n ≈ 16·σ²/δ² per variant, where δ is the smallest difference worth detecting. For a proportion p, σ² = p(1−p).
Illustrative example: baseline resolution rate 60% (σ² = 0.24). To detect a 3-point absolute lift (δ = 0.03): n ≈ 16 × 0.24 / 0.0009 ≈ 4,267 users per variant. Detecting 1 point would need about nine times more. Small effects need large samples; use a proper power calculator for real decisions.
# ab_analysis.py: difference in proportions with a 95% CI (normal approximation)
import math
def diff_ci(success_a, n_a, success_b, n_b, z=1.96):
pa, pb = success_a / n_a, success_b / n_b
se = math.sqrt(pa * (1 - pa) / n_a + pb * (1 - pb) / n_b)
d = pb - pa
return d, (d - z * se, d + z * se)
d, (lo, hi) = diff_ci(success_a=2580, n_a=4300, success_b=2710, n_b=4310) # illustrative numbers
print(f"lift {d:+.2%} 95% CI [{lo:+.2%}, {hi:+.2%}]")
Pitfalls specific to LLM experiments
- Sample ratio mismatch (SRM): if you planned 50/50 but got 52/48, assignment is broken (for example the new variant crashes and users drop out). Check SRM before reading results.
- Novelty effects: users try the new assistant more at first. Run long enough.
- Interference: agents sharing resources (rate limits, caches) across variants.
- Peeking: stopping as soon as results look good inflates false positives. Use a fixed horizon or sequential methods designed for peeking.
- Metric mismatch: a variant that answers more confidently may raise "resolution" while quietly raising later refunds or repeat contacts. Include delayed outcomes.
- Variance from non-determinism: the same variant behaves differently per request, which increases variance; plan for it.
Compliance and ethics
Tell users they are interacting with AI where laws or platform rules require it (for example obligations emerging under the EU AI Act for certain systems, and consumer-protection expectations elsewhere). Experiments must not expose one group to materially unsafe or unfair treatment; guardrails with automatic rollback are part of doing this responsibly. Check your privacy notice covers experimentation.
Worked example
A telecom in Saudi Arabia tested a new retrieval configuration for its Arabic-English support assistant. Shadow mode showed higher groundedness offline. A two-week A/B test randomized by account measured first-contact resolution (primary), with escalations, complaints and p95 latency as guardrails. Resolution improved modestly with a confidence interval excluding zero; latency p95 rose slightly but stayed within SLO. They shipped, and added the latency increase to the capacity plan.
How to measure success
Every significant LLM change goes through shadow or canary before full exposure; A/B tests have pre-registered hypotheses, primary and guardrail metrics, and SRM checks.
Video lecture: A/B tests and online experiments for LLM features
Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.
- A/B tests and online experiments
- Analogy: a new checkout system
- Staged rollout
- Design
- Sample size (rule of thumb)
- Reading results
- LLM-specific pitfalls
- Compliance and ethics
- Case: KSA telecom retrieval test
- Example: the missing 800 users (illustrative)
- Common mistakes
- Deeper: shadow mode, affordably
- Watch me do it: plan → analysis (illustrative)
- Recap
Lecture transcript
A/B tests and online experiments
Offline evals can tell you a change is safe and probably better. They cannot tell you whether real customers will resolve more issues, buy more, or come back. For that you need online experiments. In this lecture you will learn a staged rollout from shadow mode to canary to A/B test, how to design an experiment with primary and guardrail metrics, a quick way to estimate sample size, and the pitfalls that are specific to LLM features.
Analogy: a new checkout system
An analogy for staged rollouts. When a pharmacy chain introduces a new checkout system, it does not switch every branch on the same morning. It runs it quietly beside the old one in one branch, then turns it on in a few branches, compares queues and errors, and only then rolls it out everywhere, with a plan to switch back. Shadow mode, canary and A/B test are that same caution, applied to your AI feature.
Staged rollout
Roll out in stages. First, offline evals pass. Second, shadow mode: the new variant runs on live inputs in parallel, but users never see its output. You grade it with your offline graders, on the real input distribution, with zero user risk, though you do pay for the extra calls and must handle the data carefully. Third, a canary: a small share of traffic sees the new variant, with automatic rollback if guardrails break. Then the A/B test, and finally full rollout.
Design
Designing the A/B test starts with a hypothesis, like adding order-status tool calls will increase first-contact resolution without raising escalations. Randomize by user or account, not by request, so each person gets a consistent experience. Choose one primary metric that reflects value, such as resolution rate or conversion. Then choose guardrail metrics that must not get worse: safety violations, escalations, complaints, latency, cost and refunds. Write your analysis plan before you start.
Sample size (rule of thumb)
How many users do you need? A handy rule of thumb for about eighty percent power at five percent significance is sixteen times the variance, divided by the square of the smallest difference you care about, per variant. For a sixty percent resolution rate, the variance is zero point two four. To detect a three point lift, that is roughly four thousand three hundred users per variant. To detect a one point lift, about nine times more. Small effects need big samples. Use a proper calculator for real decisions.
Reading results
When results arrive, first check for sample ratio mismatch. If you planned a fifty-fifty split and got fifty-two forty-eight, something is broken, perhaps the new variant crashes and users drop out, and the results cannot be trusted. Then compute the difference with a confidence interval; the lesson includes a small function for proportions. And do not peek and stop early because it looks good. That inflates false positives unless you use methods designed for sequential monitoring.
LLM-specific pitfalls
Some pitfalls are specific to LLM features. Novelty effects: people play with a new assistant more at first, so run long enough to cover full weekly cycles. Interference: variants sharing rate limits or caches. Metric mismatch: a variant that answers more confidently can raise resolution today while quietly raising refunds or repeat contacts next week, so include delayed outcomes. And extra variance from non-determinism, since the same variant behaves differently on each request.
Compliance and ethics
Experiments with AI also carry compliance and ethical duties. Tell users they are talking to AI where laws or platform rules require it, for example obligations under the EU AI Act for certain systems and consumer protection expectations elsewhere. Do not expose one group to materially unsafe or unfair treatment; guardrails with automatic rollback are part of experimenting responsibly. And check that your privacy notice covers experimentation.
Case: KSA telecom retrieval test
A worked example from a telecom in Saudi Arabia. It tested a new retrieval configuration for its Arabic and English support assistant. Shadow mode showed better groundedness. A two-week A/B test, randomized by account, used first-contact resolution as the primary metric, with escalations, complaints and p95 latency as guardrails. Resolution improved modestly, with a confidence interval above zero. Latency rose slightly but stayed within its objective. They shipped and updated their capacity plan.
Example: the missing 800 users (illustrative)
A simple example of sample ratio mismatch. You plan a fifty-fifty split and after a week you see ten thousand users in A and nine thousand two hundred in B. That gap is far bigger than chance. Investigating, you find the new variant sometimes crashes on older Android phones before the assignment event is logged. So B is missing exactly the users with older phones, and any comparison is biased. Fix the crash, restart the test, and only then read results.
Common mistakes
Common mistakes in LLM experiments. Randomizing by request instead of by user, so the same person bounces between variants. Choosing five primary metrics and declaring victory on whichever moves. Peeking daily and stopping on the first good day. And ignoring delayed outcomes like refunds or repeat contacts. Try this now: for your next change, write down one primary metric and three guardrails before the experiment starts.
Deeper: shadow mode, affordably
One level deeper on shadow mode costs. Running the candidate in shadow doubles model calls for the shadowed traffic. So the telecom shadowed only twenty percent of conversations for one week, enough to compare groundedness on real inputs, and excluded conversations flagged as sensitive from being sent to the extra model call.
Watch me do it: plan → analysis (illustrative)
Watch me do it with the experiment plan and the analysis function. The hypothesis: adding the order-status tool increases first-contact resolution without raising escalations. Unit: randomize by account. Primary metric: resolution. Guardrails: escalations, complaints, PII detections, p95 latency and cost per conversation. Sample size: baseline resolution sixty percent, variance zero point two four, smallest lift worth detecting three points, so sixteen times zero point two four divided by zero point zero zero zero nine, about four thousand three hundred accounts per variant. At our traffic, that is about twelve days, so I round up to two full weeks to cover weekly cycles. After two weeks, first the ratio check: four thousand three hundred in A and four thousand three hundred and ten in B. Fine. Then the analysis function: two thousand five hundred and eighty resolutions in A, two thousand seven hundred and ten in B. The lift is about two point nine points, with a ninety-five percent interval from roughly one to five points, so it excludes zero. Guardrails: escalations flat, complaints flat, latency slightly up but within objective. Delayed outcome, repeat contacts within seven days: flat. Decision: ship, with a note to watch latency. All numbers illustrative.
Recap
Recap. Stage your rollout through shadow and canary before a full A/B test. Randomize by user, pick one primary metric and several guardrails, size the test in advance, check sample ratio, avoid peeking, and include delayed outcomes. Your next step: write a one-page experiment plan for your next significant LLM change, including hypothesis, metrics, sample size estimate and rollback rules.
Key takeaways
- Stage rollouts: offline evals, shadow mode, canary with auto-rollback, A/B test, full rollout.
- Randomize by user, pick one primary metric and several guardrails, and pre-register the analysis.
- Estimate sample size up front (n ≈ 16σ²/δ² per variant as a rule of thumb); small effects need large samples.
- Check sample ratio mismatch, avoid peeking, and include delayed outcomes like refunds and repeat contacts.
Try it
Write a one-page experiment plan for your next LLM change: hypothesis, unit, primary and guardrail metrics, sample-size estimate and rollback rules.