---
title: "Guardrails and human-in-the-loop | Optimize All Academy"
description: "Guardrails are layered, not single A guardrail is any control that keeps an AI system's inputs, actions or outputs within acceptable bounds. No single…"
url: https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/guardrails-and-human-in-the-loop
updated: 2026-10-05
---

Latest AI Techniques: RAG, Tool Use, Agents & MCP · Guardrails, human oversight and evaluating agents · lesson 18 of 20 · 12 min

# Guardrails and human-in-the-loop

## Guardrails are layered, not single

A **guardrail** is any control that keeps an AI system's inputs, actions or outputs within acceptable bounds. No single guardrail is sufficient; robust systems combine several layers, each catching what others miss.

## The layers

**1. Input guardrails**

- Scope checks: is the request within the product's purpose? Route out-of-scope requests politely.
- Abuse and safety classification: detect harmful requests using moderation models or classifiers.
- Personal data detection: redact or block sensitive data where it should not flow.
- Injection signals: flag suspicious instructions in documents or tool results (as one signal, not a guarantee).

**2. Action guardrails**

- Least-privilege tools and credentials.
- Allow-lists for actions, domains and recipients.
- Limits: spend caps, rate limits, maximum records changed per run.
- Approval gates for consequential actions.
- Sandboxing for code execution and browsing.

**3. Output guardrails**

- Schema validation for structured outputs.
- Policy checks: no prices quoted, no medical advice, required disclaimers present.
- Groundedness checks for RAG answers.
- Sensitive data scanning before display or sending.

**4. Process guardrails**

- Logging and audit trails.
- Monitoring and alerting.
- Incident response and kill switches to disable a feature or tool quickly.

## Designing human-in-the-loop well

Human oversight only works if it is designed for real humans:

- **Approve the right things.** Require approval where impact is high or irreversible (payments, external communications, data deletion, legal or medical outputs), not on every trivial step, or reviewers will click "approve" without reading.
- **Show what matters.** An approval screen should show the exact action, its parameters, the reason the agent gave, and what sources it relied on. "Agent wants to send an email" is not enough; show the recipient and full text.
- **Make rejection easy and informative.** Reviewers should be able to edit, reject with a reason, or escalate. Feed that back into evaluation.
- **Watch for automation bias.** People tend to over-trust confident systems. Rotate reviewers, audit samples of approved actions, and track how often reviewers change things (a near-zero rate may signal rubber-stamping, not perfection).

## Confidence-based routing

Many production systems route by risk and confidence:

```text
if action.risk == "high":                 -> require human approval
elif model_confidence < threshold
     or checks_failed:                     -> human review queue
else:                                      -> proceed, sample 5% for audit
```

"Confidence" here should come from something measurable (validation results, groundedness checks, classifier scores, agreement across samples), not simply the model saying it is sure.

## Worked example: refunds assistant

An online retailer lets an AI agent handle refund requests.

- Input: classifies the request; routes fraud signals to a human.
- Action: the agent can issue refunds only up to a set amount, only for orders belonging to the authenticated customer, only within policy windows, enforced in code.
- Above the limit or outside policy: drafts a recommendation for a human agent with the evidence.
- Output: messages checked for promises not in policy.
- Process: daily audit of a sample of automated refunds; alerts if refund volume spikes.

The model makes judgement calls; the code enforces the non-negotiables.

## Failure modes

- Guardrails only in the prompt ("never refund more than..."), which can be bypassed or ignored.
- Approval fatigue from too many low-value approvals.
- Over-blocking: guardrails so strict that legitimate users are refused, pushing them to unsafe workarounds.
- No kill switch when something goes wrong at scale.

## Hands-on: guardrails as code around a model call

The pattern below wraps an AI feature (a refund assistant) with input, action and output guardrails, and routes by risk. The model makes judgement calls; the code enforces the non-negotiables.

```python
import re
from dataclasses import dataclass

REFUND_AUTO_LIMIT = 50.00          # set by policy owners, not by the model
PII = re.compile(r"\b(\d{13,19}|\d{3}-\d{2}-\d{4})\b")      # card-like or ID-like numbers
PROMISES = re.compile(r"\b(guarantee|free forever|compensation of)\b", re.I)

@dataclass
class Decision:
    route: str       # "auto" | "review" | "block"
    reason: str

def input_guard(text: str) -> Decision | None:
    if PII.search(text):
        return Decision("review", "possible card or ID number in message; redact before AI")
    if len(text) > 8000:
        return Decision("block", "message too long")
    return None

def action_guard(action: dict, customer_id: str, orders: dict) -> Decision:
    order = orders.get(action.get("order_id"))
    if not order or order["customer_id"] != customer_id:
        return Decision("block", "order does not belong to this customer")
    if action["amount"] > order["paid"]:
        return Decision("block", "refund exceeds amount paid")
    if action["amount"] > REFUND_AUTO_LIMIT or order["days_since_delivery"] > 30:
        return Decision("review", "above auto limit or outside policy window")
    return Decision("auto", "within policy")

def output_guard(reply: str) -> Decision | None:
    if PROMISES.search(reply):
        return Decision("review", "reply contains a promise not in policy")
    return None
```

Wire these around the model call: input guard before the prompt is built, action guard before any tool executes, output guard before anything reaches the customer. Log every `Decision` with the request ID. Then measure the guardrails like classifiers: label a sample, and track how often they block legitimate requests (false positives) and miss problems (false negatives).

## Platform guardrail features

Most platforms now offer building blocks: moderation and safety classifiers, structured outputs that guarantee schema-valid JSON, tool approval settings (for example n8n's human review step for AI agent tool calls, or approval policies in agent SDKs and hosted agent platforms), spending limits, and content provenance or watermarking features for generated media. Use them, but keep your non-negotiables in your own code where you can test them, and remember that a guardrail you cannot measure is a guess.

## Going further

Measure guardrails like any classifier: how often they block legitimate requests (false positives) and how often they miss problems (false negatives), using labelled test sets including adversarial cases. Tune thresholds per risk level and review them as usage changes.

## Video lecture: Guardrails and human-in-the-loop

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Guardrails that hold
2. Analogy: airport security
3. Four layers
4. Non-negotiables live in code
5. Simple example
6. Human review that works
7. Automation bias
8. Route by risk and confidence
9. Business example (illustrative)
10. Hands-on in the lesson
11. Common mistakes
12. How you'll know it's working
13. Watch me do it: three guards
14. Recap
15. Try this now (15 minutes)

## Lecture transcript

### Guardrails that hold

A retailer tells its AI refund assistant, in the system prompt, never refund more than fifty pounds. A clever customer writes a message that talks the assistant into a two-hundred-pound refund. The prompt was a guardrail made of paper. In this lesson you'll learn why guardrails must be layered, which layers matter, how to design human review that people actually take seriously, and how to put the non-negotiables into code.

### Analogy: airport security

Here's an analogy for layered guardrails. Think of airport security. There isn't one check. There's the booking check, the ID check, the bag scanner, the body scanner, the gate check, and staff watching the whole time. Each layer misses things; together they catch almost everything. AI guardrails work the same way: input checks, action limits, output checks and monitoring, each catching what the others miss.

### Four layers

A guardrail is any control that keeps an AI system's inputs, actions or outputs within acceptable bounds. No single one is enough. Input guardrails check scope, classify abuse, detect personal data and flag injection signals. Action guardrails are least-privilege tools, allow-lists for actions and recipients, spend caps and rate limits, approval gates and sandboxes. Output guardrails validate schemas, check policy, like no prices quoted or required disclaimers present, verify groundedness and scan for sensitive data. Process guardrails are logging, monitoring, incident response and kill switches.

### Non-negotiables live in code

The key principle: enforce non-negotiables in code. Refund limits, which customer owns which order, which domains a link may point to, how many records can change in one run. The model can make judgement calls, like whether a request sounds like a genuine delivery issue, but the code decides what's allowed. A code check can't be talked around by a persuasive customer or a hidden instruction in a document.

### Simple example

A simple example of code versus prompt. Your system prompt says: never promise delivery dates. A customer pushes hard, and the model eventually writes, it will definitely arrive tomorrow. Now add a simple output check in code that looks for phrases like definitely arrive or guaranteed by, and routes those replies to a human. The prompt reduces how often it happens; the code check stops it reaching the customer when it does.

### Human review that works

Human-in-the-loop only works if it's designed for real humans. Require approval where impact is high or irreversible, like payments, external messages, deletions, legal or medical outputs, not on every trivial step, or reviewers will click approve without reading. Show what matters: the exact action, its parameters, the agent's reason and its sources. Agent wants to send an email is useless. Show the recipient and the full text. Make rejecting and editing easy, capture the reason, and feed it back into evaluation.

### Automation bias

Watch out for automation bias. People over-trust confident systems. If reviewers approve ninety-nine point eight percent of actions without a single edit, that may not mean the agent is perfect. It may mean nobody is really reading. Rotate reviewers, audit a sample of approved actions, and highlight risky elements like amounts, dates and names so attention goes where it counts.

### Route by risk and confidence

Mature systems route by risk and confidence. High-risk actions always go to a human. Low-confidence outputs or failed checks go to a review queue. Everything else proceeds, with a small sample audited. Confidence should come from something measurable, like validation results, groundedness checks, classifier scores or agreement across samples, not the model simply saying it's sure. The refund assistant in the lesson shows this: auto-refunds only within the limit and policy window for the customer's own orders, a recommendation for a human otherwise, output checks for promises not in policy, and daily sample audits.

### Business example (illustrative)

A deeper business example, illustrative. The retailer's refunds agent handles around nine hundred requests a month. About seventy percent are auto-approved within limits, twenty-five percent go to a human with a recommendation, and five percent are blocked, mostly orders not belonging to the customer. The daily sample audit found the output guard catching a few promises of compensation each week. Refund fraud attempts that used clever wording failed, because the limit and ownership checks live in code.

### Hands-on in the lesson

In the hands-on section you'll find a compact Python pattern with input, action and output guards around a refund assistant, each returning a decision to auto-approve, review or block, with a reason you can log. You'll also see which guardrail building blocks platforms now provide, from moderation classifiers and structured outputs to tool-approval steps in automation tools. Then measure your guardrails like any classifier: false positives that block legitimate customers, and false negatives that let problems through.

### Common mistakes

Common mistakes with guardrails. Guardrails that exist only in the prompt. Approval steps on everything, so reviewers stop reading. Over-blocking, so legitimate customers get refused and find workarounds. No kill switch, or one that only the original developer knows about. And never measuring guardrails, so nobody knows how many good requests they block or bad ones they miss.

### How you'll know it's working

How will you know your guardrails are working? Measure them like any classifier. On a labelled test set with adversarial cases, how many problems do they catch, and how many legitimate requests do they block? Track reviewer change rates to spot rubber-stamping. Track incidents, near misses and time to kill switch. And make sure every non-negotiable has a code check and a test proving it still works.

### Watch me do it: three guards

Watch me do it. I open the guardrails code. First, the constants: the auto-refund limit set by policy owners, a pattern for card-like numbers, and a pattern for risky promises. Next, input guard: if the message contains a card-like number, route to review for redaction; if it's too long, block. Then action guard: if the order doesn't exist or belongs to someone else, block; if the refund exceeds the amount paid, block; if it's above the limit or outside thirty days, review; otherwise auto. Output guard checks the reply for promises like guarantee or compensation. I run three cases. A fifty-pound refund on the customer's own recent order returns auto. A two-hundred-pound request returns review, above auto limit. A request for someone else's order returns block. Every decision is logged with its reason and the request ID.

### Recap

To recap: guardrails are layered across input, action, output and process. Put non-negotiables in code, design human approval for high-impact actions with the details reviewers need, watch for rubber-stamping, and route by measurable risk and confidence. Your next step: for one AI workflow, write one guardrail at each layer and mark which ones are enforced in code. Next, we'll evaluate agents properly, and learn when not to use them at all.

### Try this now (15 minutes)

Try this now. Pick one AI workflow you run or plan. Write one guardrail at each layer: input, action, output and process. Mark which are enforced in code and which only exist in the prompt. For any non-negotiable that's only in the prompt, like a refund limit or recipient allow-list, write the one-line code check that would enforce it. That list is your guardrail backlog.

## Key takeaways

- Guardrails are layered: input, action, output and process controls.
- Enforce non-negotiables in code (limits, permissions, allow-lists), not only in prompts.
- Design human approval for high-impact actions, with the details reviewers need, and watch for rubber-stamping.
- Route by measurable risk and confidence, and measure guardrails' false positives and negatives.

## Try it

For one AI workflow, list one guardrail at each layer (input, action, output, process), and mark which are enforced in code.

- [Previous: The open agent standards: A2A, Agent Skills and AGENTS.md](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/open-agent-standards)
- [Next: Securing agents: the OWASP Top 10 for Agentic Applications](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp/securing-agents-owasp-agentic)
- [All lessons of Latest AI Techniques: RAG, Tool Use, Agents & MCP](https://optimizeall.com/learn/latest-ai-techniques-rag-agents-mcp)
