Evaluating and Monitoring LLM ApplicationsMetrics and graders: code, judges and quality dimensions · Lesson 5 of 16

LLM-as-judge: design, bias and calibration

Article · 15 min · 9 min lecture

Video lecture

LLM-as-judge: design, bias and calibration

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

LLM-as-judge

  • Formats that work
  • Known biases
  • Calibrating against humans

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

When you need a judge

Many qualities cannot be checked in code: is the answer correct given the policy, is it helpful, does the tone match the brand, is the reasoning sound? LLM-as-judge uses a model to grade outputs against criteria. Done well, it scales human judgment. Done badly, it produces confident, meaningless numbers.

Judge formats

  • Pointwise with a rubric. Grade one output against explicit criteria. Prefer binary pass/fail per criterion over 1–10 scales; binary labels are easier to calibrate and more consistent.
  • Reference-guided. Provide a reference answer or key points; the judge checks consistency.
  • Pairwise. Show two outputs (for example, old vs new prompt) and ask which is better on a criterion. Good for comparing systems where absolute scoring is hard.
  • Critique then verdict. Ask the judge for a short rationale before the verdict; this often improves consistency and gives you debuggable reasons.

Known biases

Research on LLM judges (for example the "Judging LLM-as-a-Judge" MT-Bench study by Zheng and colleagues, 2023) and practitioner experience document several biases:

BiasWhat happensMitigation
Position biasIn pairwise, prefers the first (or second) answerRun both orders; count only consistent wins
Verbosity biasPrefers longer answersCriteria that reward concision; length-controlled comparisons
Self-preferencePrefers outputs from its own model familyUse a different model family as judge where feasible; calibrate
LeniencyPasses borderline casesBinary criteria with explicit fail conditions and examples
Criteria driftInterprets vague criteria inconsistentlyPrecise definitions and few-shot examples of pass and fail

Calibrating a judge against humans

A judge is a model you are deploying; evaluate it like one.

  1. Have domain experts label 100–200 outputs pass/fail on the criterion (the calibration set).
  2. Run the judge on the same outputs.
  3. Measure agreement: overall accuracy, true positive rate and true negative rate separately (a judge that passes everything looks accurate if most outputs are good), and a chance-corrected statistic such as Cohen's kappa.
  4. Inspect disagreements; refine the rubric and examples; re-measure on a held-out split.
  5. Recalibrate when you change the judge model or prompt.

Hands-on: a binary rubric judge

# judge.py: binary rubric judge with rationale (Anthropic SDK; adapt for other providers)
import json, os
from anthropic import Anthropic

client = Anthropic()
JUDGE_MODEL = os.environ["JUDGE_MODEL"]  # current model ID from your provider's docs

RUBRIC = """You are grading a customer-support answer for POLICY CORRECTNESS only.
PASS if every policy statement in the answer is consistent with the policy excerpt
for the customer's country. FAIL if any statement contradicts or goes beyond the policy
(e.g. promises a refund window not in the policy), or if it applies another country's policy.
Ignore tone and length.
Examples:
- Policy AE: 7-day returns. Answer: "You can return within 7 days." -> PASS
- Policy AE: 7-day returns. Answer: "You have 14 days." -> FAIL (wrong window)
Return JSON: {"rationale": "<one or two sentences>", "verdict": "PASS" | "FAIL"}"""

def judge(policy: str, country: str, question: str, answer: str) -> dict:
    content = f"{RUBRIC}\n\nCountry: {country}\nPolicy excerpt:\n{policy}\n\nQuestion:\n{question}\n\nAnswer:\n{answer}"
    try:
        msg = client.messages.create(model=JUDGE_MODEL, max_tokens=300, temperature=0,
                                     messages=[{"role": "user", "content": content}])
        text = msg.content[0].text
        data = json.loads(text[text.find("{"): text.rfind("}") + 1])
        assert data["verdict"] in ("PASS", "FAIL")
        return data
    except Exception as e:
        return {"verdict": "ERROR", "rationale": f"judge error: {e}"}  # count errors separately
# calibrate.py: agreement with human labels
from sklearn.metrics import cohen_kappa_score, confusion_matrix
human = [...]  # "PASS"/"FAIL" from experts
model = [...]  # judge verdicts on the same items (exclude ERROR, report them separately)
tn, fp, fn, tp = confusion_matrix(human, model, labels=["FAIL", "PASS"]).ravel()
print("TPR (passes caught):", tp / (tp + fn), " TNR (fails caught):", tn / (tn + fp))
print("Cohen's kappa:", cohen_kappa_score(human, model))

Many providers also support structured outputs or JSON modes that guarantee parseable responses; use them where available.

Pairwise comparison done right

Compare Answer A and Answer B for the question below on CONCISENESS WITHOUT LOSING
REQUIRED INFORMATION. Required points: {points}. Output "A", "B" or "TIE" with one sentence why.

Run each pair twice with A and B swapped. Count a win only if both runs agree; otherwise record a tie. Report win rate with a confidence interval.

Worked example

A UK insurance comparison site used a 1–10 "helpfulness" judge; scores hovered around 8 regardless of prompt changes. Switching to four binary criteria (correct cover explanation, mentions exclusions, no advice beyond regulated scope, plain English) calibrated on 150 expert labels exposed that 12% of answers strayed into personal recommendations, a compliance concern the 1–10 score had hidden.

Pitfalls

  • Vague criteria ("is it good?").
  • Uncalibrated judges reported as truth.
  • One judge for many criteria in one prompt; split criteria for clarity.
  • Ignoring judge errors and cost; log them.

How to measure success

Each judge has a documented calibration (TPR, TNR, kappa) on a held-out human-labeled set, and is recalibrated when changed.

Key takeaways

  • Prefer binary pass/fail per criterion, one criterion per judge prompt, with pass and fail examples.
  • Mitigate position, verbosity, self-preference and leniency biases; swap order in pairwise comparisons.
  • Calibrate judges against 100–200 expert labels, reporting TPR, TNR and Cohen's kappa.
  • Recalibrate whenever the judge model or prompt changes; log judge errors separately.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A pairwise judge prefers Answer A 70% of the time when A is shown first and 40% when shown second. What bias is this?
  2. Most outputs in your set are good. Your judge shows 90% accuracy vs human labels. What else must you check?

Put it into practice

Collect 100 expert pass/fail labels for one criterion, run the judge from this lesson, and report TPR, TNR and kappa with three rubric improvements.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.