---
title: "Human review workflows and feedback loops"
description: "Humans remain the ground truth Every metric in this course ultimately rests on human judgment: judges are calibrated against human labels, datasets are…"
url: https://optimizeall.com/learn/llm-evals-and-observability/human-review-workflows
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Production observability, monitoring and experimentation · lesson 14 of 16 · 13 min

# Human review workflows and feedback loops

## Humans remain the ground truth

Every metric in this course ultimately rests on human judgment: judges are calibrated against human labels, datasets are verified by experts, taxonomies come from people reading traces. A **human review workflow** makes that judgment systematic, scalable and auditable, instead of ad hoc screenshots in a chat channel.

## What to send to human review

You cannot review everything. Build review queues from several streams:

| Stream | Why |
|---|---|
| Random sample | Unbiased estimate of quality; finds silent failures |
| User-flagged (thumbs down, complaints) | High yield of real problems |
| Automated flags (judge FAIL, low confidence, PII detector, safety classifier) | Verify automated signals; catch false positives |
| High-stakes routes (refunds, medical, legal, financial) | Risk-weighted oversight |
| Disagreements (judge vs code check, or two judges) | Improve graders |
| New slices (new language, new product) | Early warning |

## Designing the review task

- **Clear rubric** with the same criteria as your judges (binary per criterion, with examples).
- **Full context on one screen:** user input, retrieved sources, tool calls, output, customer country or tier.
- **Structured labels plus a free-text note** on the first failure (feeding error analysis).
- **Corrected output** where useful: an edited ideal answer becomes a golden item.
- **Time budget:** most items should take a minute or two; if longer, the rubric or interface needs work.

## Inter-annotator agreement

Before trusting labels, have two reviewers label the same subset (for example 10–20%) and measure agreement: percent agreement plus a chance-corrected statistic such as Cohen's kappa (two raters) or Krippendorff's alpha (more raters, missing data). Low agreement means the rubric is ambiguous, not that reviewers are careless. Hold calibration sessions, discuss disagreements, update the guidelines with the examples.

## Closing the loop

A review workflow is only valuable if labels flow back:

1. Reviewed failures → **error-analysis taxonomy** counts.
2. Confirmed failures with corrected answers → **golden and regression datasets** (versioned).
3. Reviewer labels on judged items → **judge calibration** updates.
4. Patterns → **product fixes** (content, prompts, retrieval, tools, policy).
5. Summary → **weekly quality review** with owners and actions.

## Operational concerns

- **Privacy and access:** reviewers see real customer data; restrict access by role, redact where possible, log access, and follow your data-protection obligations. Consider data residency when reviewers are in different countries from customers.
- **Reviewer wellbeing:** reviewing harmful content (abuse, self-harm) is taxing; rotate duties, provide support, blur by default where possible.
- **Expertise:** some criteria need domain experts (clinicians, lawyers, finance specialists); others can be handled by trained generalists. Route accordingly.
- **Incentives:** measure reviewers on agreement and quality, not speed alone.
- **Regulatory oversight:** for higher-risk uses, documented human oversight may be expected by regulators or customers (for example under the EU AI Act for high-risk systems, and in sector guidance from financial regulators). Keep audit trails.

## Hands-on: a minimal review queue schema

```sql
CREATE TABLE review_items (
  id            BIGSERIAL PRIMARY KEY,
  trace_id      TEXT NOT NULL,
  stream        TEXT NOT NULL,           -- random | user_flag | judge_fail | high_stakes | disagreement | new_slice
  route         TEXT, country TEXT, language TEXT,
  assigned_to   TEXT,
  status        TEXT NOT NULL DEFAULT 'open',   -- open | done | escalated
  created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE review_labels (
  item_id       BIGINT REFERENCES review_items(id),
  reviewer      TEXT NOT NULL,
  criterion     TEXT NOT NULL,           -- policy_correct | grounded | tone | privacy ...
  verdict       TEXT NOT NULL CHECK (verdict IN ('PASS','FAIL','NA')),
  first_failure_note TEXT,
  corrected_output   TEXT,
  labeled_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);
```

Most observability platforms offer annotation queues with similar concepts; the schema shows what you need regardless of tool.

## Worked example

A fintech in Dubai routed all conversations mentioning disputes or chargebacks to a compliance reviewer queue, plus 2% random traffic to trained generalists. A monthly calibration session raised agreement noticeably after rubric examples were added. Corrected answers from reviewers grew the golden set by dozens of items a month, and the policy judge was recalibrated quarterly against reviewer labels.

## Pitfalls

- **Reviewing only complaints.**
- **Labels that never leave the tool.**
- **No agreement measurement.**
- **Unrestricted access to customer conversations.**

## How to measure success

A steady weekly review volume across streams, agreement above your threshold, and a visible pipeline from labels to datasets, judges and fixes.

## Video lecture: Human review workflows and feedback loops

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Human review workflows
2. Analogy: the standards desk
3. Review streams
4. The review task
5. Inter-annotator agreement
6. Close the loop
7. Operations
8. Oversight and audit
9. Case: Dubai fintech
10. Example: 34 of 40 agree (illustrative)
11. Common mistakes
12. Scenario: one month of review (illustrative)
13. Deeper: protecting reviewers
14. Watch me do it: one week of review
15. Recap

## Lecture transcript

### Human review workflows

Every metric in this course rests, in the end, on human judgment. Judges are calibrated against human labels. Datasets are verified by experts. Taxonomies come from people reading traces. In this lecture you will learn to make that human judgment systematic: what to send for review, how to design the review task, how to measure whether reviewers agree, and how to close the loop so labels actually improve your system.

### Analogy: the standards desk

An analogy for human review. Think of a newspaper's standards desk. Reporters and automated tools produce a lot of copy. The standards editors do not read everything, but they read a smart sample, every complaint, and every high-stakes story, using a clear style guide, and they meet regularly to agree on tricky calls. Their corrections change the style guide and train the reporters. Your review queue is that standards desk for your AI system.

### Review streams

You cannot review everything, so build queues from several streams. A random sample gives an unbiased quality estimate and finds silent failures. User-flagged items, like thumbs down and complaints, have a high yield of real problems. Automated flags, like judge failures or PII detections, need verification. High-stakes routes, such as refunds, medical or financial questions, deserve risk-weighted oversight. Disagreements between graders improve the graders. And new slices, like a newly launched language, give early warning.

### The review task

Design the review task carefully. Use the same binary criteria as your judges, with examples. Put full context on one screen: the user's input, retrieved sources, tool calls, the output, and customer country or tier. Ask for structured labels plus a short note on the first failure, which feeds your error analysis. Where useful, have reviewers write a corrected answer, which becomes a golden item. And aim for one or two minutes per item. If it takes longer, fix the rubric or the interface.

### Inter-annotator agreement

Before trusting labels, measure agreement. Have two reviewers label the same ten to twenty percent of items. Report percent agreement plus a chance-corrected statistic: Cohen's kappa for two raters, or Krippendorff's alpha for more raters or missing labels. Low agreement usually means the rubric is ambiguous, not that people are careless. Hold calibration sessions, discuss the disagreements, and add the resolved examples to your guidelines.

### Close the loop

Now the part most teams miss: closing the loop. Reviewed failures update your error-analysis taxonomy counts. Confirmed failures with corrected answers go into versioned golden and regression datasets. Reviewer labels on judged items update judge calibration. Patterns become product fixes: content, prompts, retrieval, tools or policy. And a weekly quality review assigns owners and actions. Labels that never leave the review tool are wasted effort.

### Operations

Operational concerns matter too. Reviewers see real customer data, so restrict access by role, redact where possible, log access, and consider data residency when reviewers sit in a different country from customers. Protect reviewer wellbeing when content includes abuse or self-harm: rotate duties, blur by default and provide support. Route criteria needing clinicians, lawyers or finance specialists to experts. And measure reviewers on agreement and quality, not speed alone.

### Oversight and audit

For higher-risk uses, documented human oversight may be expected by regulators or customers, for example under the EU AI Act for high-risk systems, or in financial regulators' guidance. Keep an audit trail of who reviewed what and when. The lesson includes a minimal database schema for a review queue and labels, which shows the concepts you need no matter which annotation tool you choose.

### Case: Dubai fintech

A worked example. A fintech in Dubai routed every conversation mentioning disputes or chargebacks to a compliance reviewer queue, plus two percent of random traffic to trained generalists. Monthly calibration sessions raised agreement after they added rubric examples. Reviewers' corrected answers grew the golden set every month, and the policy judge was recalibrated quarterly against their labels. Human review became the engine of improvement, not a cost center.

### Example: 34 of 40 agree (illustrative)

A simple example of measuring agreement. Two reviewers label the same forty answers for policy correctness. They agree on thirty-four, which sounds good. But most answers are passes, so some agreement would happen by chance. Cohen's kappa corrects for that, and here it comes out only moderate. Looking at the six disagreements, five involve answers that promise a callback within a time the policy does not guarantee. The rubric never mentioned callbacks. Add one example, re-label, and agreement improves.

### Common mistakes

Common mistakes in human review. Reviewing only complaints, which hides silent failures. Letting labels pile up inside the review tool without flowing into datasets or judge calibration. Never measuring agreement, so nobody knows if the labels are reliable. And giving every reviewer access to every customer conversation. Quick question: what happened to the last hundred labels your reviewers produced? Where are they now?

### Scenario: one month of review (illustrative)

And a realistic scenario with illustrative numbers. An insurance broker in Dubai routes two percent of random chats plus every conversation mentioning claims to a review queue. In the first month, reviewers label around six hundred items. Forty confirmed failures become new golden cases, two recurring issues become prompt and retrieval fixes, and one reviewer note reveals that the bot was quoting an expired promotion. Because the labels flowed back into datasets and judges, the next month's automated scores reflected those lessons without any extra manual work.

### Deeper: protecting reviewers

One level deeper on reviewer wellbeing. On the abuse-report route, reviewers see content blurred by default and click to reveal, rotate off that queue every two weeks, and have a clear way to skip an item without penalty. Quality of labels improved, because tired, distressed reviewers make inconsistent judgments.

### Watch me do it: one week of review

Watch me do it with the review queue schema from the lesson. Each review item has a trace ID, a stream, route, country and language, an assignee and a status. Each label row has the reviewer, a criterion, a verdict of pass, fail or not applicable, a first-failure note and an optional corrected output. First, I fill the queue for one week: two percent random, every thumbs-down, every judge failure, every conversation on the refunds route, and every Arabic conversation, because we launched Arabic last month. That is about three hundred items. Second, I assign twenty of them to two reviewers each. Third, after the week I compute agreement on those twenty for policy correctness: they agree on sixteen, and kappa comes out moderate. The four disagreements all involve partial refunds. I write one rubric example for partial refunds and read it out in a fifteen-minute calibration call. Fourth, I query the labels table for confirmed failures with corrected outputs: thirty-one rows. Those become new golden items, with the reviewer as source. Fifth, I export verdicts on judge-failed items to update the judge's calibration set. Sixth, the weekly quality review takes the top three failure notes and assigns an owner to each. The schema is simple, and every field has a job downstream.

### Recap

Recap. Build review queues from random, flagged, automated, high-stakes, disagreement and new-slice streams. Design fast, full-context tasks with the same rubric as your judges. Measure agreement and calibrate. Close the loop into taxonomies, datasets, judges and fixes, and run it with privacy, wellbeing and audit in mind. Your next step: set up a review queue with at least three streams and double-label twenty items to measure agreement this week.

## Key takeaways

- Build review queues from random, user-flagged, automated, high-stakes, disagreement and new-slice streams.
- Give reviewers full context, the judges' rubric, and space for first-failure notes and corrected answers.
- Measure inter-annotator agreement (kappa or alpha) and calibrate; low agreement signals an ambiguous rubric.
- Close the loop into taxonomies, datasets, judge calibration and fixes, with privacy, wellbeing and audit trails.

## Try it

Set up a review queue with at least three streams, double-label 20 items, compute agreement, and export confirmed failures into your dataset.

- [Previous: A/B tests and online experiments for LLM features](https://optimizeall.com/learn/llm-evals-and-observability/ab-tests-and-online-experiments)
- [Next: Capstone part 1: dataset, interface and harness](https://optimizeall.com/learn/llm-evals-and-observability/capstone-build-the-harness)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
