Latest AI Techniques: RAG, Tool Use, Agents & MCPGuardrails, human oversight and evaluating agents · Lesson 18 of 20

Guardrails and human-in-the-loop

Article · 12 min · 8 min lecture

Video lecture

Guardrails and human-in-the-loop

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Guardrails that hold

  • Layered, not single
  • Input, action, output, process
  • Human review that works

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Guardrails are layered, not single

A guardrail is any control that keeps an AI system's inputs, actions or outputs within acceptable bounds. No single guardrail is sufficient; robust systems combine several layers, each catching what others miss.

The layers

1. Input guardrails

  • Scope checks: is the request within the product's purpose? Route out-of-scope requests politely.
  • Abuse and safety classification: detect harmful requests using moderation models or classifiers.
  • Personal data detection: redact or block sensitive data where it should not flow.
  • Injection signals: flag suspicious instructions in documents or tool results (as one signal, not a guarantee).

2. Action guardrails

  • Least-privilege tools and credentials.
  • Allow-lists for actions, domains and recipients.
  • Limits: spend caps, rate limits, maximum records changed per run.
  • Approval gates for consequential actions.
  • Sandboxing for code execution and browsing.

3. Output guardrails

  • Schema validation for structured outputs.
  • Policy checks: no prices quoted, no medical advice, required disclaimers present.
  • Groundedness checks for RAG answers.
  • Sensitive data scanning before display or sending.

4. Process guardrails

  • Logging and audit trails.
  • Monitoring and alerting.
  • Incident response and kill switches to disable a feature or tool quickly.

Designing human-in-the-loop well

Human oversight only works if it is designed for real humans:

  • Approve the right things. Require approval where impact is high or irreversible (payments, external communications, data deletion, legal or medical outputs), not on every trivial step, or reviewers will click "approve" without reading.
  • Show what matters. An approval screen should show the exact action, its parameters, the reason the agent gave, and what sources it relied on. "Agent wants to send an email" is not enough; show the recipient and full text.
  • Make rejection easy and informative. Reviewers should be able to edit, reject with a reason, or escalate. Feed that back into evaluation.
  • Watch for automation bias. People tend to over-trust confident systems. Rotate reviewers, audit samples of approved actions, and track how often reviewers change things (a near-zero rate may signal rubber-stamping, not perfection).

Confidence-based routing

Many production systems route by risk and confidence:

if action.risk == "high":                 -> require human approval
elif model_confidence < threshold
     or checks_failed:                     -> human review queue
else:                                      -> proceed, sample 5% for audit

"Confidence" here should come from something measurable (validation results, groundedness checks, classifier scores, agreement across samples), not simply the model saying it is sure.

Worked example: refunds assistant

An online retailer lets an AI agent handle refund requests.

  • Input: classifies the request; routes fraud signals to a human.
  • Action: the agent can issue refunds only up to a set amount, only for orders belonging to the authenticated customer, only within policy windows, enforced in code.
  • Above the limit or outside policy: drafts a recommendation for a human agent with the evidence.
  • Output: messages checked for promises not in policy.
  • Process: daily audit of a sample of automated refunds; alerts if refund volume spikes.

The model makes judgement calls; the code enforces the non-negotiables.

Failure modes

  • Guardrails only in the prompt ("never refund more than..."), which can be bypassed or ignored.
  • Approval fatigue from too many low-value approvals.
  • Over-blocking: guardrails so strict that legitimate users are refused, pushing them to unsafe workarounds.
  • No kill switch when something goes wrong at scale.

Hands-on: guardrails as code around a model call

The pattern below wraps an AI feature (a refund assistant) with input, action and output guardrails, and routes by risk. The model makes judgement calls; the code enforces the non-negotiables.

import re
from dataclasses import dataclass

REFUND_AUTO_LIMIT = 50.00          # set by policy owners, not by the model
PII = re.compile(r"\b(\d{13,19}|\d{3}-\d{2}-\d{4})\b")      # card-like or ID-like numbers
PROMISES = re.compile(r"\b(guarantee|free forever|compensation of)\b", re.I)

@dataclass
class Decision:
    route: str       # "auto" | "review" | "block"
    reason: str

def input_guard(text: str) -> Decision | None:
    if PII.search(text):
        return Decision("review", "possible card or ID number in message; redact before AI")
    if len(text) > 8000:
        return Decision("block", "message too long")
    return None

def action_guard(action: dict, customer_id: str, orders: dict) -> Decision:
    order = orders.get(action.get("order_id"))
    if not order or order["customer_id"] != customer_id:
        return Decision("block", "order does not belong to this customer")
    if action["amount"] > order["paid"]:
        return Decision("block", "refund exceeds amount paid")
    if action["amount"] > REFUND_AUTO_LIMIT or order["days_since_delivery"] > 30:
        return Decision("review", "above auto limit or outside policy window")
    return Decision("auto", "within policy")

def output_guard(reply: str) -> Decision | None:
    if PROMISES.search(reply):
        return Decision("review", "reply contains a promise not in policy")
    return None

Wire these around the model call: input guard before the prompt is built, action guard before any tool executes, output guard before anything reaches the customer. Log every Decision with the request ID. Then measure the guardrails like classifiers: label a sample, and track how often they block legitimate requests (false positives) and miss problems (false negatives).

Platform guardrail features

Most platforms now offer building blocks: moderation and safety classifiers, structured outputs that guarantee schema-valid JSON, tool approval settings (for example n8n's human review step for AI agent tool calls, or approval policies in agent SDKs and hosted agent platforms), spending limits, and content provenance or watermarking features for generated media. Use them, but keep your non-negotiables in your own code where you can test them, and remember that a guardrail you cannot measure is a guess.

Going further

Measure guardrails like any classifier: how often they block legitimate requests (false positives) and how often they miss problems (false negatives), using labelled test sets including adversarial cases. Tune thresholds per risk level and review them as usage changes.

Key takeaways

  • Guardrails are layered: input, action, output and process controls.
  • Enforce non-negotiables in code (limits, permissions, allow-lists), not only in prompts.
  • Design human approval for high-impact actions, with the details reviewers need, and watch for rubber-stamping.
  • Route by measurable risk and confidence, and measure guardrails' false positives and negatives.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which guardrail is most robust for a refund limit?
  2. Reviewers approve 99.8% of agent actions without edits. What might this indicate?
  3. What should a good approval screen show?

Put it into practice

For one AI workflow, list one guardrail at each layer (input, action, output, process), and mark which are enforced in code.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.