Evaluating and Monitoring LLM ApplicationsProduction observability, monitoring and experimentation · Lesson 14 of 16
Human review workflows and feedback loops
Video lecture
Human review workflows and feedback loops
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Human review workflows
Every metric in this course rests, in the end, on human judgment. Judges are calibrated against human labels. Datasets are verified by experts. Taxonomies come from people reading traces. In this lecture you will learn to make that human judgment systematic: what to send for review, how to design the review task, how to measure whether reviewers agree, and how to close the loop so labels actually improve your system.
0:30 Analogy: the standards desk
An analogy for human review. Think of a newspaper's standards desk. Reporters and automated tools produce a lot of copy. The standards editors do not read everything, but they read a smart sample, every complaint, and every high-stakes story, using a clear style guide, and they meet regularly to agree on tricky calls. Their corrections change the style guide and train the reporters. Your review queue is that standards desk for your AI system.
1:02 Review streams
You cannot review everything, so build queues from several streams. A random sample gives an unbiased quality estimate and finds silent failures. User-flagged items, like thumbs down and complaints, have a high yield of real problems. Automated flags, like judge failures or PII detections, need verification. High-stakes routes, such as refunds, medical or financial questions, deserve risk-weighted oversight. Disagreements between graders improve the graders. And new slices, like a newly launched language, give early warning.
1:35 The review task
Design the review task carefully. Use the same binary criteria as your judges, with examples. Put full context on one screen: the user's input, retrieved sources, tool calls, the output, and customer country or tier. Ask for structured labels plus a short note on the first failure, which feeds your error analysis. Where useful, have reviewers write a corrected answer, which becomes a golden item. And aim for one or two minutes per item. If it takes longer, fix the rubric or the interface.
2:12 Inter-annotator agreement
Before trusting labels, measure agreement. Have two reviewers label the same ten to twenty percent of items. Report percent agreement plus a chance-corrected statistic: Cohen's kappa for two raters, or Krippendorff's alpha for more raters or missing labels. Low agreement usually means the rubric is ambiguous, not that people are careless. Hold calibration sessions, discuss the disagreements, and add the resolved examples to your guidelines.
2:40 Close the loop
Now the part most teams miss: closing the loop. Reviewed failures update your error-analysis taxonomy counts. Confirmed failures with corrected answers go into versioned golden and regression datasets. Reviewer labels on judged items update judge calibration. Patterns become product fixes: content, prompts, retrieval, tools or policy. And a weekly quality review assigns owners and actions. Labels that never leave the review tool are wasted effort.
3:08 Operations
Operational concerns matter too. Reviewers see real customer data, so restrict access by role, redact where possible, log access, and consider data residency when reviewers sit in a different country from customers. Protect reviewer wellbeing when content includes abuse or self-harm: rotate duties, blur by default and provide support. Route criteria needing clinicians, lawyers or finance specialists to experts. And measure reviewers on agreement and quality, not speed alone.
3:38 Oversight and audit
For higher-risk uses, documented human oversight may be expected by regulators or customers, for example under the EU AI Act for high-risk systems, or in financial regulators' guidance. Keep an audit trail of who reviewed what and when. The lesson includes a minimal database schema for a review queue and labels, which shows the concepts you need no matter which annotation tool you choose.
4:06 Case: Dubai fintech
A worked example. A fintech in Dubai routed every conversation mentioning disputes or chargebacks to a compliance reviewer queue, plus two percent of random traffic to trained generalists. Monthly calibration sessions raised agreement after they added rubric examples. Reviewers' corrected answers grew the golden set every month, and the policy judge was recalibrated quarterly against their labels. Human review became the engine of improvement, not a cost center.
4:36 Example: 34 of 40 agree (illustrative)
A simple example of measuring agreement. Two reviewers label the same forty answers for policy correctness. They agree on thirty-four, which sounds good. But most answers are passes, so some agreement would happen by chance. Cohen's kappa corrects for that, and here it comes out only moderate. Looking at the six disagreements, five involve answers that promise a callback within a time the policy does not guarantee. The rubric never mentioned callbacks. Add one example, re-label, and agreement improves.
5:10 Common mistakes
Common mistakes in human review. Reviewing only complaints, which hides silent failures. Letting labels pile up inside the review tool without flowing into datasets or judge calibration. Never measuring agreement, so nobody knows if the labels are reliable. And giving every reviewer access to every customer conversation. Quick question: what happened to the last hundred labels your reviewers produced? Where are they now?
5:37 Scenario: one month of review (illustrative)
And a realistic scenario with illustrative numbers. An insurance broker in Dubai routes two percent of random chats plus every conversation mentioning claims to a review queue. In the first month, reviewers label around six hundred items. Forty confirmed failures become new golden cases, two recurring issues become prompt and retrieval fixes, and one reviewer note reveals that the bot was quoting an expired promotion. Because the labels flowed back into datasets and judges, the next month's automated scores reflected those lessons without any extra manual work.
6:15 Deeper: protecting reviewers
One level deeper on reviewer wellbeing. On the abuse-report route, reviewers see content blurred by default and click to reveal, rotate off that queue every two weeks, and have a clear way to skip an item without penalty. Quality of labels improved, because tired, distressed reviewers make inconsistent judgments.
6:36 Watch me do it: one week of review
Watch me do it with the review queue schema from the lesson. Each review item has a trace ID, a stream, route, country and language, an assignee and a status. Each label row has the reviewer, a criterion, a verdict of pass, fail or not applicable, a first-failure note and an optional corrected output. First, I fill the queue for one week: two percent random, every thumbs-down, every judge failure, every conversation on the refunds route, and every Arabic conversation, because we launched Arabic last month. That is about three hundred items. Second, I assign twenty of them to two reviewers each. Third, after the week I compute agreement on those twenty for policy correctness: they agree on sixteen, and kappa comes out moderate. The four disagreements all involve partial refunds. I write one rubric example for partial refunds and read it out in a fifteen-minute calibration call. Fourth, I query the labels table for confirmed failures with corrected outputs: thirty-one rows. Those become new golden items, with the reviewer as source. Fifth, I export verdicts on judge-failed items to update the judge's calibration set. Sixth, the weekly quality review takes the top three failure notes and assigns an owner to each. The schema is simple, and every field has a job downstream.
8:09 Recap
Recap. Build review queues from random, flagged, automated, high-stakes, disagreement and new-slice streams. Design fast, full-context tasks with the same rubric as your judges. Measure agreement and calibrate. Close the loop into taxonomies, datasets, judges and fixes, and run it with privacy, wellbeing and audit in mind. Your next step: set up a review queue with at least three streams and double-label twenty items to measure agreement this week.
Humans remain the ground truth
Every metric in this course ultimately rests on human judgment: judges are calibrated against human labels, datasets are verified by experts, taxonomies come from people reading traces. A human review workflow makes that judgment systematic, scalable and auditable, instead of ad hoc screenshots in a chat channel.
What to send to human review
You cannot review everything. Build review queues from several streams:
| Stream | Why |
|---|---|
| Random sample | Unbiased estimate of quality; finds silent failures |
| User-flagged (thumbs down, complaints) | High yield of real problems |
| Automated flags (judge FAIL, low confidence, PII detector, safety classifier) | Verify automated signals; catch false positives |
| High-stakes routes (refunds, medical, legal, financial) | Risk-weighted oversight |
| Disagreements (judge vs code check, or two judges) | Improve graders |
| New slices (new language, new product) | Early warning |
Designing the review task
- Clear rubric with the same criteria as your judges (binary per criterion, with examples).
- Full context on one screen: user input, retrieved sources, tool calls, output, customer country or tier.
- Structured labels plus a free-text note on the first failure (feeding error analysis).
- Corrected output where useful: an edited ideal answer becomes a golden item.
- Time budget: most items should take a minute or two; if longer, the rubric or interface needs work.
Inter-annotator agreement
Before trusting labels, have two reviewers label the same subset (for example 10–20%) and measure agreement: percent agreement plus a chance-corrected statistic such as Cohen's kappa (two raters) or Krippendorff's alpha (more raters, missing data). Low agreement means the rubric is ambiguous, not that reviewers are careless. Hold calibration sessions, discuss disagreements, update the guidelines with the examples.
Closing the loop
A review workflow is only valuable if labels flow back:
- Reviewed failures → error-analysis taxonomy counts.
- Confirmed failures with corrected answers → golden and regression datasets (versioned).
- Reviewer labels on judged items → judge calibration updates.
- Patterns → product fixes (content, prompts, retrieval, tools, policy).
- Summary → weekly quality review with owners and actions.
Operational concerns
- Privacy and access: reviewers see real customer data; restrict access by role, redact where possible, log access, and follow your data-protection obligations. Consider data residency when reviewers are in different countries from customers.
- Reviewer wellbeing: reviewing harmful content (abuse, self-harm) is taxing; rotate duties, provide support, blur by default where possible.
- Expertise: some criteria need domain experts (clinicians, lawyers, finance specialists); others can be handled by trained generalists. Route accordingly.
- Incentives: measure reviewers on agreement and quality, not speed alone.
- Regulatory oversight: for higher-risk uses, documented human oversight may be expected by regulators or customers (for example under the EU AI Act for high-risk systems, and in sector guidance from financial regulators). Keep audit trails.
Hands-on: a minimal review queue schema
CREATE TABLE review_items (
id BIGSERIAL PRIMARY KEY,
trace_id TEXT NOT NULL,
stream TEXT NOT NULL, -- random | user_flag | judge_fail | high_stakes | disagreement | new_slice
route TEXT, country TEXT, language TEXT,
assigned_to TEXT,
status TEXT NOT NULL DEFAULT 'open', -- open | done | escalated
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE review_labels (
item_id BIGINT REFERENCES review_items(id),
reviewer TEXT NOT NULL,
criterion TEXT NOT NULL, -- policy_correct | grounded | tone | privacy ...
verdict TEXT NOT NULL CHECK (verdict IN ('PASS','FAIL','NA')),
first_failure_note TEXT,
corrected_output TEXT,
labeled_at TIMESTAMPTZ NOT NULL DEFAULT now()
);Most observability platforms offer annotation queues with similar concepts; the schema shows what you need regardless of tool.
Worked example
A fintech in Dubai routed all conversations mentioning disputes or chargebacks to a compliance reviewer queue, plus 2% random traffic to trained generalists. A monthly calibration session raised agreement noticeably after rubric examples were added. Corrected answers from reviewers grew the golden set by dozens of items a month, and the policy judge was recalibrated quarterly against reviewer labels.
Pitfalls
- Reviewing only complaints.
- Labels that never leave the tool.
- No agreement measurement.
- Unrestricted access to customer conversations.
How to measure success
A steady weekly review volume across streams, agreement above your threshold, and a visible pipeline from labels to datasets, judges and fixes.
Key takeaways
- Build review queues from random, user-flagged, automated, high-stakes, disagreement and new-slice streams.
- Give reviewers full context, the judges' rubric, and space for first-failure notes and corrected answers.
- Measure inter-annotator agreement (kappa or alpha) and calibrate; low agreement signals an ambiguous rubric.
- Close the loop into taxonomies, datasets, judge calibration and fixes, with privacy, wellbeing and audit trails.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Set up a review queue with at least three streams, double-label 20 items, compute agreement, and export confirmed failures into your dataset.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.