Advanced Prompt EngineeringEvaluation sets and grading · Lesson 13 of 17

Grading outputs: code checks, rubrics and LLM-as-judge

Article · 13 min · 7 min lecture

Video lecture

Grading outputs: code checks, rubrics and LLM-as-judge

12 chapters · about 7 min · full transcript

Coming soon

Chapter 1 of 12

Grading outputs

  • Code checks, rubrics, LLM judges
  • Writing anchored rubrics
  • Judge biases
  • Calibration

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Choosing the right grader

Once you have test cases, you need a way to score outputs. Use the cheapest grader that measures what you care about reliably:

  1. Code-based grading. Exact match, enum match, regular expressions, schema validation, numeric tolerance, keyword presence or absence. Fast, cheap, perfectly consistent. Use it wherever the criterion is objective.
  2. Model-based grading (LLM-as-judge). A model scores outputs against a rubric. Scales to subjective criteria like helpfulness, tone and completeness.
  3. Human grading. The ground truth for subjective quality, and essential for calibrating the other two. Slow and expensive, so use it strategically.

Most mature setups combine them: code gates first, model grading for nuanced dimensions, and periodic human review to keep the model grader honest.

Writing rubrics

A rubric turns taste into criteria. Strong rubrics are specific, observable and anchored:

Criterion: Answers the customer's actual question
5 = Direct answer in the first two sentences, fully correct per policy
3 = Answer present but buried, or partially complete
1 = Does not answer, or answer contradicts policy

Criterion: Tone
5 = Calm, empathetic, no blame, no jargon
3 = Neutral but stiff, or one jargon term
1 = Defensive, blaming or sarcastic

Tips:

  • Prefer binary or small scales (pass/fail, or 1-3-5 with anchors). Wide scales like 1 to 10 invite inconsistency, from humans and models alike.
  • Score one criterion at a time, ideally in separate judge calls for important criteria, so a strong tone does not inflate a weak answer.
  • Ask for evidence before the score: "Quote the sentence that answers the question, then score."

LLM-as-judge: powerful, with known caveats

Research and practice have identified biases that you should design around:

  • Position bias. In pairwise comparisons ("Is A or B better?"), judges can favour the first or second position. Mitigate by running both orders and only counting consistent wins.
  • Verbosity bias. Longer answers can be rated higher even when not better. Include length expectations in the rubric.
  • Self-preference. A model may favour outputs that resemble its own style. Where possible use a judge from a different model family, or validate against humans.
  • Leniency and inconsistency. Judges may give high scores broadly, or vary between runs. Use anchored rubrics, low randomness settings where available, and repeated sampling on important decisions.
  • Blind spots. A judge cannot reliably verify facts it does not know. Give it the reference material or grading notes.

A judge prompt template

You are grading a support reply against a rubric. Be strict.

<policy>{{policy}}</policy>
<customer_message>{{input}}</customer_message>
<reply>{{output}}</reply>
<grading_notes>{{notes}}</grading_notes>

Criterion: The reply is fully supported by the policy.
First list each factual claim in the reply and whether the policy supports it.
Then output JSON: {"claims": [...], "verdict": "PASS" | "FAIL"}

Calibrate the judge against humans

Before trusting a judge, have humans grade a sample (for example 50 outputs) using the same rubric. Compare: how often does the judge agree with the humans? Where does it disagree, and in which direction? Adjust the rubric or judge prompt until agreement is acceptable for your stakes, and re-check periodically, especially after changing the judge model.

Worked example

A marketing team compares two prompts for product descriptions. The judge prefers prompt B in most comparisons. Suspicious, they check: B's outputs are much longer. After adding "length within the 80 to 120 word target is required; longer is not better" and swapping positions, the preference largely disappears, and humans prefer A for clarity. Without the calibration step, they would have shipped the worse prompt.

Hands-on: a calibrated judge

import json, os, anthropic
client = anthropic.Anthropic()
JUDGE_MODEL = os.environ.get("JUDGE_MODEL", "claude-opus-5")  # consider a different family from the generator

VERDICT = {"type": "object", "properties": {
    "claims": {"type": "array", "items": {"type": "object", "properties": {
        "claim": {"type": "string"}, "supported": {"type": "boolean"}},
        "required": ["claim", "supported"], "additionalProperties": False}},
    "verdict": {"type": "string", "enum": ["PASS", "FAIL"]}},
    "required": ["claims", "verdict"], "additionalProperties": False}

def judge(policy, message, reply):
    resp = client.messages.create(
        model=JUDGE_MODEL, max_tokens=2048,
        output_config={"format": {"type": "json_schema", "schema": VERDICT}},
        messages=[{"role": "user", "content": f"""Grade strictly.
<policy>{policy}</policy><customer_message>{message}</customer_message><reply>{reply}</reply>
Criterion: every factual claim in the reply is supported by the policy.
List each claim with supported true/false, then the verdict (FAIL if any claim is unsupported)."""}],
    )
    return json.loads(next(b.text for b in resp.content if b.type == "text"))["verdict"]

def agreement(samples):
    """samples: list of dicts with policy, message, reply and human ('PASS'/'FAIL')."""
    tp = fp = tn = fn = 0
    for s in samples:
        j, h = judge(s["policy"], s["message"], s["reply"]), s["human"]
        tp += j == h == "FAIL"; tn += j == h == "PASS"
        fp += j == "FAIL" and h == "PASS"; fn += j == "PASS" and h == "FAIL"
    n = len(samples)
    print(f"agreement={(tp + tn) / n:.0%}  missed_failures={fn}  false_alarms={fp}")

Look beyond raw agreement: for a safety criterion, missed failures (judge says PASS, human says FAIL) matter far more than false alarms. Read every disagreement and adjust the rubric, then re-run on a fresh sample so you do not overfit the judge to the calibration set.

Pairwise comparisons done right

When comparing prompt A with prompt B, show the judge both outputs twice with the order swapped, and count a win only when both orders agree. Report ties separately. This simple step removes much of the position bias.

Reasoning judges

Judges using a reasoning model or higher effort are often more accurate on complex rubrics (multi-claim factuality, code review) but cost more per grade. Run your calibration at two settings and choose the cheaper one that meets your agreement target.

Going further

Track grader metrics as first-class: agreement rate with humans, pass rates by tag, and judge cost. Store every judged output with the judge's reasoning so you can audit surprising scores. When the stakes are high, report results with uncertainty (for example the range across repeated runs) rather than a single number.

Key takeaways

  • Use the cheapest reliable grader: code for objective checks, LLM judges for nuance, humans to calibrate.
  • Good rubrics are specific, observable and anchored, with small scales and one criterion per judgement.
  • LLM judges show position, verbosity and self-preference biases; design around them.
  • Calibrate judges against human grades before trusting them, and re-check after changes.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. In a pairwise LLM-as-judge comparison, how do you reduce position bias?
  2. Which criterion should be graded with code rather than a model?
  3. A judge consistently prefers longer answers that humans rate worse. What should you do first?

Put it into practice

Write an anchored rubric (3 criteria, 1-3-5 scale) for one output type. Grade 20 outputs yourself, then with an LLM judge, and compute how often you agree.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.