Evaluating and Monitoring LLM ApplicationsProduction observability, monitoring and experimentation · Lesson 14 of 16

Human review workflows and feedback loops

Article · 13 min · 9 min lecture

Video lecture

Human review workflows and feedback loops

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Human review workflows

  • What to review
  • Designing the task
  • Agreement
  • Closing the loop

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Humans remain the ground truth

Every metric in this course ultimately rests on human judgment: judges are calibrated against human labels, datasets are verified by experts, taxonomies come from people reading traces. A human review workflow makes that judgment systematic, scalable and auditable, instead of ad hoc screenshots in a chat channel.

What to send to human review

You cannot review everything. Build review queues from several streams:

StreamWhy
Random sampleUnbiased estimate of quality; finds silent failures
User-flagged (thumbs down, complaints)High yield of real problems
Automated flags (judge FAIL, low confidence, PII detector, safety classifier)Verify automated signals; catch false positives
High-stakes routes (refunds, medical, legal, financial)Risk-weighted oversight
Disagreements (judge vs code check, or two judges)Improve graders
New slices (new language, new product)Early warning

Designing the review task

  • Clear rubric with the same criteria as your judges (binary per criterion, with examples).
  • Full context on one screen: user input, retrieved sources, tool calls, output, customer country or tier.
  • Structured labels plus a free-text note on the first failure (feeding error analysis).
  • Corrected output where useful: an edited ideal answer becomes a golden item.
  • Time budget: most items should take a minute or two; if longer, the rubric or interface needs work.

Inter-annotator agreement

Before trusting labels, have two reviewers label the same subset (for example 10–20%) and measure agreement: percent agreement plus a chance-corrected statistic such as Cohen's kappa (two raters) or Krippendorff's alpha (more raters, missing data). Low agreement means the rubric is ambiguous, not that reviewers are careless. Hold calibration sessions, discuss disagreements, update the guidelines with the examples.

Closing the loop

A review workflow is only valuable if labels flow back:

  1. Reviewed failures → error-analysis taxonomy counts.
  2. Confirmed failures with corrected answers → golden and regression datasets (versioned).
  3. Reviewer labels on judged items → judge calibration updates.
  4. Patterns → product fixes (content, prompts, retrieval, tools, policy).
  5. Summary → weekly quality review with owners and actions.

Operational concerns

  • Privacy and access: reviewers see real customer data; restrict access by role, redact where possible, log access, and follow your data-protection obligations. Consider data residency when reviewers are in different countries from customers.
  • Reviewer wellbeing: reviewing harmful content (abuse, self-harm) is taxing; rotate duties, provide support, blur by default where possible.
  • Expertise: some criteria need domain experts (clinicians, lawyers, finance specialists); others can be handled by trained generalists. Route accordingly.
  • Incentives: measure reviewers on agreement and quality, not speed alone.
  • Regulatory oversight: for higher-risk uses, documented human oversight may be expected by regulators or customers (for example under the EU AI Act for high-risk systems, and in sector guidance from financial regulators). Keep audit trails.

Hands-on: a minimal review queue schema

CREATE TABLE review_items (
  id            BIGSERIAL PRIMARY KEY,
  trace_id      TEXT NOT NULL,
  stream        TEXT NOT NULL,           -- random | user_flag | judge_fail | high_stakes | disagreement | new_slice
  route         TEXT, country TEXT, language TEXT,
  assigned_to   TEXT,
  status        TEXT NOT NULL DEFAULT 'open',   -- open | done | escalated
  created_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE TABLE review_labels (
  item_id       BIGINT REFERENCES review_items(id),
  reviewer      TEXT NOT NULL,
  criterion     TEXT NOT NULL,           -- policy_correct | grounded | tone | privacy ...
  verdict       TEXT NOT NULL CHECK (verdict IN ('PASS','FAIL','NA')),
  first_failure_note TEXT,
  corrected_output   TEXT,
  labeled_at    TIMESTAMPTZ NOT NULL DEFAULT now()
);

Most observability platforms offer annotation queues with similar concepts; the schema shows what you need regardless of tool.

Worked example

A fintech in Dubai routed all conversations mentioning disputes or chargebacks to a compliance reviewer queue, plus 2% random traffic to trained generalists. A monthly calibration session raised agreement noticeably after rubric examples were added. Corrected answers from reviewers grew the golden set by dozens of items a month, and the policy judge was recalibrated quarterly against reviewer labels.

Pitfalls

  • Reviewing only complaints.
  • Labels that never leave the tool.
  • No agreement measurement.
  • Unrestricted access to customer conversations.

How to measure success

A steady weekly review volume across streams, agreement above your threshold, and a visible pipeline from labels to datasets, judges and fixes.

Key takeaways

  • Build review queues from random, user-flagged, automated, high-stakes, disagreement and new-slice streams.
  • Give reviewers full context, the judges' rubric, and space for first-failure notes and corrected answers.
  • Measure inter-annotator agreement (kappa or alpha) and calibrate; low agreement signals an ambiguous rubric.
  • Close the loop into taxonomies, datasets, judge calibration and fixes, with privacy, wellbeing and audit trails.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Two reviewers agree on only 60% of labels for 'policy correct'. What is the most likely fix?
  2. Which destination should reviewer-corrected answers feed?

Put it into practice

Set up a review queue with at least three streams, double-label 20 items, compute agreement, and export confirmed failures into your dataset.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.