Advanced Prompt EngineeringEvaluation sets and grading · Lesson 13 of 17
Grading outputs: code checks, rubrics and LLM-as-judge
Video lecture
Grading outputs: code checks, rubrics and LLM-as-judge
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Grading outputs
A marketing team asked an AI judge to compare two prompts for product descriptions. The judge clearly preferred prompt B. They almost shipped it. Then someone noticed that B's outputs were simply longer. In this lecture you will learn to choose the right grader, write rubrics that produce consistent scores, design around the known biases of LLM judges, and calibrate a judge against human grades before you trust it.
0:30 Why grading matters
Why does grading deserve a whole lecture? Because your evaluation is only as good as the thing that scores it. A sloppy grader can make a worse prompt look better, and you will ship it with confidence. Think of a grader like a referee. If the referee is biased toward one team, the scoreboard lies, however well the match is played. So before you trust any score, you need to know how your grader behaves, and where it goes wrong.
1:05 Choose the grader
Use the cheapest grader that measures what you care about reliably. Code-based grading, exact match, enum match, regular expressions, schema validation, numeric tolerance or keyword checks, is fast, cheap and perfectly consistent; use it wherever the criterion is objective. Model-based grading, LLM-as-judge, scores outputs against a rubric and scales to subjective criteria like helpfulness, tone and completeness. And human grading is the ground truth for subjective quality and essential for calibrating the other two. Mature setups combine all three.
1:39 Anchored rubrics
A rubric turns taste into criteria. Strong rubrics are specific, observable and anchored. For answers the customer's actual question: five means a direct answer in the first two sentences, fully correct per policy; three means the answer is present but buried or partial; one means it does not answer, or contradicts policy. Prefer binary or small scales like one, three, five with anchors; wide scales like one to ten invite inconsistency from humans and models alike. Score one criterion at a time, ideally in separate judge calls for the important ones. And ask for evidence before the score.
2:22 Judge biases
LLM judges are powerful, with known biases. Position bias: in pairwise comparisons, judges can favour the first or second answer; run both orders and count only consistent wins. Verbosity bias: longer answers can score higher without being better, so put length expectations in the rubric. Self-preference: a model may favour outputs resembling its own style; consider a judge from a different model family. Leniency and inconsistency: use anchored rubrics and repeated sampling on important decisions. And blind spots: a judge cannot verify facts it does not know, so give it the reference material or grading notes.
3:04 A judge prompt that works
Here is a judge design from the lesson. The judge receives the policy, the customer message, the reply and the grading notes. For the criterion, the reply is fully supported by the policy, it first lists each factual claim and whether the policy supports it, then outputs a verdict, pass or fail, in a schema-enforced JSON object. Listing claims first grounds the verdict, and structured output makes it easy to aggregate. A reasoning model or higher effort often judges complex rubrics more accurately, at higher cost; test two settings and choose the cheaper one that meets your agreement target.
3:47 Calibrate against humans
Never trust a judge until you have calibrated it. Have humans grade a sample, say fifty outputs, with the same rubric. Compare. How often does the judge agree? Where does it disagree, and in which direction? For a safety criterion, missed failures, where the judge says pass and the human says fail, matter far more than false alarms. Read every disagreement, adjust the rubric or judge prompt, and re-check on a fresh sample, so you do not overfit the judge to the calibration set. Re-calibrate whenever you change the judge model.
4:27 Worked example
Back to the marketing team. After adding a rule, length within the eighty to one hundred and twenty word target is required, longer is not better, and swapping the order of the two outputs, the judge's preference for B largely disappeared. Human reviewers preferred A for clarity. Without calibration, they would have shipped the worse prompt. When comparing prompts pairwise, always run both orders, count a win only when both agree, and report ties separately.
5:00 Operate your graders
Treat grader metrics as first-class. Track agreement with humans, pass rates by tag, and judge cost. Store every judged output with the judge's reasoning, so you can audit surprising scores. When the stakes are high, report results with uncertainty, like the range across repeated runs, rather than a single number that looks more precise than it is.
5:25 Example 1: a politeness rubric
A simple worked example. You want to grade whether customer replies are polite. Your first rubric: rate politeness from one to ten. Two humans score the same reply a four and an eight. Now anchor it. Pass: greets the customer, no blame, no sarcasm, thanks them or acknowledges the issue. Fail: blames the customer, uses sarcasm, or dismisses the issue. Add: quote the sentence that most affects your decision, then give pass or fail. The two humans now agree on nearly every reply, and so does the LLM judge. Small, observable, evidence first.
6:05 Example 2: quiz explanations (illustrative)
Now a business scenario, with illustrative numbers. An e-learning company grades AI-written quiz explanations for about two thousand questions a quarter. They use an LLM judge with a rubric on correctness, clarity and whether the explanation teaches why distractors are wrong. Before trusting it, two subject experts grade eighty explanations. Initial judge agreement with the experts: seventy percent, and worse, the judge passed eleven explanations the experts failed, mostly subtle factual errors. The team gives the judge the reference answer and the source material, asks it to list each claim and whether the source supports it before scoring, and switches to a judge from a different model family than the generator. Agreement rises to around ninety-one percent, with two missed failures in eighty. They re-check on a fresh sample of fifty and see similar results. Illustrative figures.
7:05 Recap
To recap. Use code for objective checks, LLM judges for nuance, and humans for ground truth. Write anchored rubrics with small scales, one criterion per judgement, and evidence before scores. Design around position, verbosity and self-preference bias, and calibrate judges against human grades before trusting them. Try this now: write a three-criterion rubric on a one-three-five scale, grade twenty outputs yourself, grade them with an LLM judge, and compute agreement. Next: automated prompt optimisation.
Choosing the right grader
Once you have test cases, you need a way to score outputs. Use the cheapest grader that measures what you care about reliably:
- Code-based grading. Exact match, enum match, regular expressions, schema validation, numeric tolerance, keyword presence or absence. Fast, cheap, perfectly consistent. Use it wherever the criterion is objective.
- Model-based grading (LLM-as-judge). A model scores outputs against a rubric. Scales to subjective criteria like helpfulness, tone and completeness.
- Human grading. The ground truth for subjective quality, and essential for calibrating the other two. Slow and expensive, so use it strategically.
Most mature setups combine them: code gates first, model grading for nuanced dimensions, and periodic human review to keep the model grader honest.
Writing rubrics
A rubric turns taste into criteria. Strong rubrics are specific, observable and anchored:
Criterion: Answers the customer's actual question
5 = Direct answer in the first two sentences, fully correct per policy
3 = Answer present but buried, or partially complete
1 = Does not answer, or answer contradicts policy
Criterion: Tone
5 = Calm, empathetic, no blame, no jargon
3 = Neutral but stiff, or one jargon term
1 = Defensive, blaming or sarcasticTips:
- Prefer binary or small scales (pass/fail, or 1-3-5 with anchors). Wide scales like 1 to 10 invite inconsistency, from humans and models alike.
- Score one criterion at a time, ideally in separate judge calls for important criteria, so a strong tone does not inflate a weak answer.
- Ask for evidence before the score: "Quote the sentence that answers the question, then score."
LLM-as-judge: powerful, with known caveats
Research and practice have identified biases that you should design around:
- Position bias. In pairwise comparisons ("Is A or B better?"), judges can favour the first or second position. Mitigate by running both orders and only counting consistent wins.
- Verbosity bias. Longer answers can be rated higher even when not better. Include length expectations in the rubric.
- Self-preference. A model may favour outputs that resemble its own style. Where possible use a judge from a different model family, or validate against humans.
- Leniency and inconsistency. Judges may give high scores broadly, or vary between runs. Use anchored rubrics, low randomness settings where available, and repeated sampling on important decisions.
- Blind spots. A judge cannot reliably verify facts it does not know. Give it the reference material or grading notes.
A judge prompt template
You are grading a support reply against a rubric. Be strict.
<policy>{{policy}}</policy>
<customer_message>{{input}}</customer_message>
<reply>{{output}}</reply>
<grading_notes>{{notes}}</grading_notes>
Criterion: The reply is fully supported by the policy.
First list each factual claim in the reply and whether the policy supports it.
Then output JSON: {"claims": [...], "verdict": "PASS" | "FAIL"}Calibrate the judge against humans
Before trusting a judge, have humans grade a sample (for example 50 outputs) using the same rubric. Compare: how often does the judge agree with the humans? Where does it disagree, and in which direction? Adjust the rubric or judge prompt until agreement is acceptable for your stakes, and re-check periodically, especially after changing the judge model.
Worked example
A marketing team compares two prompts for product descriptions. The judge prefers prompt B in most comparisons. Suspicious, they check: B's outputs are much longer. After adding "length within the 80 to 120 word target is required; longer is not better" and swapping positions, the preference largely disappears, and humans prefer A for clarity. Without the calibration step, they would have shipped the worse prompt.
Hands-on: a calibrated judge
import json, os, anthropic
client = anthropic.Anthropic()
JUDGE_MODEL = os.environ.get("JUDGE_MODEL", "claude-opus-5") # consider a different family from the generator
VERDICT = {"type": "object", "properties": {
"claims": {"type": "array", "items": {"type": "object", "properties": {
"claim": {"type": "string"}, "supported": {"type": "boolean"}},
"required": ["claim", "supported"], "additionalProperties": False}},
"verdict": {"type": "string", "enum": ["PASS", "FAIL"]}},
"required": ["claims", "verdict"], "additionalProperties": False}
def judge(policy, message, reply):
resp = client.messages.create(
model=JUDGE_MODEL, max_tokens=2048,
output_config={"format": {"type": "json_schema", "schema": VERDICT}},
messages=[{"role": "user", "content": f"""Grade strictly.
<policy>{policy}</policy><customer_message>{message}</customer_message><reply>{reply}</reply>
Criterion: every factual claim in the reply is supported by the policy.
List each claim with supported true/false, then the verdict (FAIL if any claim is unsupported)."""}],
)
return json.loads(next(b.text for b in resp.content if b.type == "text"))["verdict"]
def agreement(samples):
"""samples: list of dicts with policy, message, reply and human ('PASS'/'FAIL')."""
tp = fp = tn = fn = 0
for s in samples:
j, h = judge(s["policy"], s["message"], s["reply"]), s["human"]
tp += j == h == "FAIL"; tn += j == h == "PASS"
fp += j == "FAIL" and h == "PASS"; fn += j == "PASS" and h == "FAIL"
n = len(samples)
print(f"agreement={(tp + tn) / n:.0%} missed_failures={fn} false_alarms={fp}")Look beyond raw agreement: for a safety criterion, missed failures (judge says PASS, human says FAIL) matter far more than false alarms. Read every disagreement and adjust the rubric, then re-run on a fresh sample so you do not overfit the judge to the calibration set.
Pairwise comparisons done right
When comparing prompt A with prompt B, show the judge both outputs twice with the order swapped, and count a win only when both orders agree. Report ties separately. This simple step removes much of the position bias.
Reasoning judges
Judges using a reasoning model or higher effort are often more accurate on complex rubrics (multi-claim factuality, code review) but cost more per grade. Run your calibration at two settings and choose the cheaper one that meets your agreement target.
Going further
Track grader metrics as first-class: agreement rate with humans, pass rates by tag, and judge cost. Store every judged output with the judge's reasoning so you can audit surprising scores. When the stakes are high, report results with uncertainty (for example the range across repeated runs) rather than a single number.
Key takeaways
- Use the cheapest reliable grader: code for objective checks, LLM judges for nuance, humans to calibrate.
- Good rubrics are specific, observable and anchored, with small scales and one criterion per judgement.
- LLM judges show position, verbosity and self-preference biases; design around them.
- Calibrate judges against human grades before trusting them, and re-check after changes.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Write an anchored rubric (3 criteria, 1-3-5 scale) for one output type. Grade 20 outputs yourself, then with an LLM judge, and compute how often you agree.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.