Skip to content

Advanced Prompt Engineering · Decomposition, reasoning and verification · lesson 7 of 17 · 11 min

Self-critique and verification loops

Generate, check, revise

Humans rarely produce their best work in one pass; neither do models. A verification loop adds a checking stage between generation and delivery. The key insight is that checking is often easier than generating: it is easier to verify that every figure in a summary appears in the source than to write the summary perfectly in one go.

Three kinds of checks

  1. Deterministic checks (code). Schema validation, word limits, required sections present, numbers matching source, links resolving, banned terms absent. Cheap, fast and reliable; always do these first.
  2. Model-based checks. A second prompt (or the same model in a separate call) assesses the output against explicit criteria: factual support, tone, policy compliance, completeness.
  3. Human checks. For high-stakes outputs, a person reviews, ideally guided by the model's flagged concerns so their time goes where it matters.

Designing a good critique prompt

Vague critique ("make it better") produces cosmetic edits. Effective critique is criteria-based and evidence-seeking:

You are reviewing a draft customer email against the source policy.

<policy>{{policy}}</policy>
<draft>{{draft}}</draft>

For each criterion, answer PASS or FAIL with a one-line reason:
1. Every factual claim about refunds is supported by the policy (quote it).
2. No promise is made that the policy does not allow.
3. The customer's actual question is answered in the first two sentences.
4. Tone is calm and non-defensive.

Then output a JSON object: {"all_pass": bool, "fixes": [..]}

A separate revise step then applies only the listed fixes. Separating critique from revision makes the loop inspectable and stops the reviser from rewriting things that were fine.

Same model or different model?

Using the same model to check its own work can work well for surface issues but can share its blind spots, for example repeating the same factual mistake. Mitigations:

  • Give the checker different information: the source documents, the rules, or a tool (search, calculator, code execution) the generator did not use.
  • Use a different prompt framing: the checker's job is to find problems, not to confirm.
  • For high-stakes work, use a different model or a human for the final check.

Worked example: financial summary

An analyst uses AI to summarise quarterly results for a client newsletter.

  1. Generate the summary with quotes and figures tagged with source line numbers.
  2. Code check: extract every number in the summary and confirm it appears in the source (after normalising formats).
  3. Model check: "Does any sentence imply causation the source does not state?"
  4. Revise with the flagged fixes.
  5. Human spot-check of anything the model check flagged.

The loop catches the most common failure: a correct number attached to the wrong period.

How many rounds?

Usually one critique-and-revise round captures most of the gain. Further rounds give diminishing returns and can make text blander as the model sands off anything distinctive. Put a hard cap (for example two rounds), and escalate to a human when the checker still fails.

The over-verification trap

Verification costs tokens and time. Some recent models self-check effectively without being asked, and piling on "double-check everything three times" instructions can slow them down without raising accuracy. Add verification where your evaluation shows errors, not everywhere by habit.

Failure modes

  • Rubber-stamping. The checker says PASS to everything. Test it with deliberately flawed drafts; a checker that cannot catch planted errors is decoration.
  • Nitpicking. The checker always finds something, so revisions never end. Require evidence and severity, and ignore low-severity items.
  • Checker drift. You improve the generator but not the checker's criteria, so it checks for yesterday's problems.

Hands-on: a critique-then-revise loop with a planted-error test

import json, os, anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")

CRITIQUE_SCHEMA = {
    "type": "object",
    "properties": {
        "criteria": {"type": "array", "items": {
            "type": "object",
            "properties": {"id": {"type": "integer"}, "evidence": {"type": "string"},
                           "verdict": {"type": "string", "enum": ["PASS", "FAIL"]},
                           "severity": {"type": "string", "enum": ["low", "high"]}},
            "required": ["id", "evidence", "verdict", "severity"],
            "additionalProperties": False}},
        "fixes": {"type": "array", "items": {"type": "string"}},
    },
    "required": ["criteria", "fixes"], "additionalProperties": False,
}

def critique(policy, draft):
    resp = client.messages.create(
        model=MODEL, max_tokens=4096,
        output_config={"format": {"type": "json_schema", "schema": CRITIQUE_SCHEMA}},
        messages=[{"role": "user", "content": f"""Review the draft against the policy.
<policy>{policy}</policy>
<draft>{draft}</draft>
Criteria:
1. Every refund claim is supported by the policy (quote it as evidence).
2. No promise the policy does not allow.
3. The customer's question is answered in the first two sentences.
4. Calm, non-defensive tone.
List fixes only for high-severity failures."""}],
    )
    return json.loads(next(b.text for b in resp.content if b.type == "text"))

def revise(draft, fixes):
    resp = client.messages.create(
        model=MODEL, max_tokens=2048,
        messages=[{"role": "user", "content": "Apply only these fixes; change nothing else.\n"
                   f"<fixes>{json.dumps(fixes)}</fixes>\n<draft>{draft}</draft>"}],
    )
    return "".join(b.text for b in resp.content if b.type == "text")

def verify(policy, draft, max_rounds=2):
    for _ in range(max_rounds):
        report = critique(policy, draft)
        if not report["fixes"]:
            return draft, "pass"
        draft = revise(draft, report["fixes"])
    return draft, "escalate_to_human"

Now test the checker itself. Create ten drafts with a known, planted error (a refund promise the policy forbids, a wrong day count) and ten clean drafts. Run critique on all twenty and compute recall (planted errors caught) and false-positive rate (clean drafts flagged). If recall is low, strengthen the criteria or give the checker better reference material; if false positives are high, require evidence and severity.

Verification with tools beats verification by opinion

Where a check can be executed, execute it: run generated code against tests, recompute totals with code, check that quotes exist verbatim in sources, validate URLs. Some platforms let the model run code in a sandbox during its answer; a critique step that can calculate is far stronger than one that can only read.

Going further

Measure your checker like any classifier: seed a set of outputs with known errors and compute how many it catches (recall) and how many clean outputs it wrongly flags (false positives). This turns "we have a review step" into "our review step catches most planted factual errors at an acceptable false-alarm rate", which is a claim you can defend.

Video lecture: Self-critique and verification loops

Lecture coming soon · 12 chapters · about 8 minutes. Read the full transcript below.

  1. Self-critique and verification loops
  2. Why verify?
  3. Three kinds of checks
  4. Critique prompts that work
  5. Same model or different?
  6. Worked example: financial summary
  7. Cap the rounds
  8. Verify where it matters
  9. Test the checker
  10. Example 1: a five-bullet summary
  11. Example 2: client portfolio updates (illustrative)
  12. Recap

Lecture transcript

Self-critique and verification loops

Humans rarely produce their best work in one pass, and neither do models. But the fix is not simply asking the model to try harder. In this lecture you will learn to build verification loops: the three kinds of checks, how to write a critique prompt that finds real problems, whether the same model should check its own work, how many revision rounds are worth it, and how to test your checker so it is not just decoration.

Why verify?

Why does this matter? Because in production, one confident error can undo a hundred good outputs. A wrong figure in a client report, or a promise your policy does not allow, costs trust that takes months to rebuild. Verification is how you turn a model that is usually right into a system that is reliably right. Think of it like a newsroom: reporters write fast, and editors check the facts before anything is printed. Your pipeline needs an editor too.

Three kinds of checks

The key insight is that checking is often easier than generating. It is easier to verify that every figure in a summary appears in the source than to write the perfect summary in one go. A verification loop exploits this by adding a checking stage between generation and delivery. And there are three kinds of checks. Deterministic checks in code: schema validity, word limits, required sections, numbers matching the source, banned terms absent. Model-based checks: a separate call assessing factual support, tone, policy compliance or completeness. And human checks for high-stakes outputs, ideally guided by the model's flagged concerns.

Critique prompts that work

Vague critique, like make it better, produces cosmetic edits. Effective critique is criteria-based and evidence-seeking. For a customer email, for example: every factual claim about refunds is supported by the policy, quote it. No promise is made that the policy does not allow. The customer's question is answered in the first two sentences. Tone is calm and non-defensive. For each, answer pass or fail with evidence. Then a separate revise step applies only the listed fixes. Separating critique from revision makes the loop inspectable, and stops the reviser from rewriting things that were fine.

Same model or different?

Should the same model check its own work? It works well for surface issues, but it can share its own blind spots, repeating the same factual mistake. Three mitigations help. Give the checker different information, such as the source documents, the rules, or a tool like search, a calculator or code execution that the generator did not use. Use a different framing: the checker's job is to find problems, not to confirm. And for high-stakes work, use a different model or a human for the final check.

Worked example: financial summary

Here is a verification loop for a financial summary in a client newsletter. Generate the summary with figures tagged with source line numbers. Code check: extract every number and confirm it appears in the source after normalising formats. Model check: does any sentence imply causation the source does not state? Revise with the flagged fixes. Human spot-check anything the model flagged. This loop catches the most common failure: a correct number attached to the wrong period.

Cap the rounds

How many rounds? Usually one critique-and-revise round captures most of the gain. More rounds give diminishing returns and can make text blander, as the model sands off anything distinctive. Set a hard cap, for example two rounds, and escalate to a human if the checker still fails. The lesson's hands-on code does exactly this: a critique function returns structured verdicts and fixes, a revise function applies only those fixes, and after two rounds the draft is marked for human review.

Verify where it matters

Watch the over-verification trap. Some recent models self-check effectively without being asked. Piling on double-check everything three times slows them down without raising accuracy. Add verification where your evaluation shows errors, not everywhere by habit. And where a check can be executed, execute it. Run generated code against tests, recompute totals, confirm quotes exist verbatim. Some platforms let the model run code in a sandbox during its answer, and a critique step that can calculate is far stronger than one that can only read.

Test the checker

Finally, test the checker itself. Three failure modes: rubber-stamping, where the checker passes everything; nitpicking, where it always finds something and revisions never end; and checker drift, where criteria still target yesterday's problems. The fix is to measure it like any classifier. Plant known errors in ten drafts and leave ten drafts clean. Run the checker on all twenty and compute recall, the planted errors it caught, and the false-positive rate, the clean drafts it wrongly flagged. Now you can say our review step catches most planted errors at an acceptable false-alarm rate, and defend it.

Example 1: a five-bullet summary

A simple worked example. You ask a model to summarise a two-page article in five bullets, and to include every statistic mentioned. It writes the bullets. Now a code check: extract every number from the summary and confirm each one appears in the article. One number, sixty-two percent, does not appear; the article said sixty-two thousand. Then a model check with a single criterion: does any bullet claim something the article does not state? It flags one bullet that turned a correlation into a cause. Two quick checks, two errors caught, before anyone reads the summary.

Example 2: client portfolio updates (illustrative)

Now a business scenario, with illustrative numbers. A financial advisory firm in Dubai sends about one hundred and twenty personalised portfolio updates to clients every month, drafted by AI from the portfolio data. Compliance requires that no update implies guaranteed returns and that every figure matches the statement. They add a verification loop. Code checks every figure against the portfolio system. A critique step, using a different model family, checks three criteria: no guarantee language, no advice outside the client's risk profile, and no unsupported causal claims. Failures get one revision, then go to a human. To trust the checker, they plant errors in forty test drafts; it catches thirty-seven, and wrongly flags three of forty clean drafts. That is good enough, with humans reviewing every flagged draft. Illustrative figures.

Recap

To recap. Checking is easier than generating, so add a verification stage. Run code checks first, then criteria-based model checks with evidence, then humans for high stakes. Give checkers different information or tools, cap revision rounds, avoid habitual over-verification, and measure your checker with planted errors. Try this now: write a four-criterion critique prompt for an output you produce regularly, plant five errors in sample drafts, and measure how many the checker catches. Next module: prompting tools and agents.

Key takeaways

  • Checking is often easier than generating, so add a verification stage before delivery.
  • Run cheap deterministic code checks first, then criteria-based model checks, then humans for high stakes.
  • Give the checker different information or tools to avoid sharing the generator's blind spots.
  • Cap revision rounds, test the checker with planted errors, and avoid habitual over-verification.

Try it

Write a 4-criterion critique prompt for an output you produce regularly. Plant 5 errors in sample drafts and measure how many the checker catches.