---
title: "LLM-as-judge: design, bias and calibration"
description: "When you need a judge Many qualities cannot be checked in code: is the answer correct given the policy, is it helpful, does the tone match the brand, is…"
url: https://optimizeall.com/learn/llm-evals-and-observability/llm-as-judge
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Metrics and graders: code, judges and quality dimensions · lesson 5 of 16 · 15 min

# LLM-as-judge: design, bias and calibration

## When you need a judge

Many qualities cannot be checked in code: is the answer correct given the policy, is it helpful, does the tone match the brand, is the reasoning sound? **LLM-as-judge** uses a model to grade outputs against criteria. Done well, it scales human judgment. Done badly, it produces confident, meaningless numbers.

## Judge formats

- **Pointwise with a rubric.** Grade one output against explicit criteria. Prefer **binary pass/fail per criterion** over 1–10 scales; binary labels are easier to calibrate and more consistent.
- **Reference-guided.** Provide a reference answer or key points; the judge checks consistency.
- **Pairwise.** Show two outputs (for example, old vs new prompt) and ask which is better on a criterion. Good for comparing systems where absolute scoring is hard.
- **Critique then verdict.** Ask the judge for a short rationale before the verdict; this often improves consistency and gives you debuggable reasons.

## Known biases

Research on LLM judges (for example the "Judging LLM-as-a-Judge" MT-Bench study by Zheng and colleagues, 2023) and practitioner experience document several biases:

| Bias | What happens | Mitigation |
|---|---|---|
| Position bias | In pairwise, prefers the first (or second) answer | Run both orders; count only consistent wins |
| Verbosity bias | Prefers longer answers | Criteria that reward concision; length-controlled comparisons |
| Self-preference | Prefers outputs from its own model family | Use a different model family as judge where feasible; calibrate |
| Leniency | Passes borderline cases | Binary criteria with explicit fail conditions and examples |
| Criteria drift | Interprets vague criteria inconsistently | Precise definitions and few-shot examples of pass and fail |

## Calibrating a judge against humans

A judge is a model you are deploying; evaluate it like one.

1. Have domain experts label 100–200 outputs pass/fail on the criterion (the calibration set).
2. Run the judge on the same outputs.
3. Measure agreement: overall accuracy, **true positive rate and true negative rate** separately (a judge that passes everything looks accurate if most outputs are good), and a chance-corrected statistic such as **Cohen's kappa**.
4. Inspect disagreements; refine the rubric and examples; re-measure on a held-out split.
5. Recalibrate when you change the judge model or prompt.

## Hands-on: a binary rubric judge

```python
# judge.py: binary rubric judge with rationale (Anthropic SDK; adapt for other providers)
import json, os
from anthropic import Anthropic

client = Anthropic()
JUDGE_MODEL = os.environ["JUDGE_MODEL"]  # current model ID from your provider's docs

RUBRIC = """You are grading a customer-support answer for POLICY CORRECTNESS only.
PASS if every policy statement in the answer is consistent with the policy excerpt
for the customer's country. FAIL if any statement contradicts or goes beyond the policy
(e.g. promises a refund window not in the policy), or if it applies another country's policy.
Ignore tone and length.
Examples:
- Policy AE: 7-day returns. Answer: "You can return within 7 days." -> PASS
- Policy AE: 7-day returns. Answer: "You have 14 days." -> FAIL (wrong window)
Return JSON: {"rationale": "<one or two sentences>", "verdict": "PASS" | "FAIL"}"""

def judge(policy: str, country: str, question: str, answer: str) -> dict:
    content = f"{RUBRIC}\n\nCountry: {country}\nPolicy excerpt:\n{policy}\n\nQuestion:\n{question}\n\nAnswer:\n{answer}"
    try:
        msg = client.messages.create(model=JUDGE_MODEL, max_tokens=300, temperature=0,
                                     messages=[{"role": "user", "content": content}])
        text = msg.content[0].text
        data = json.loads(text[text.find("{"): text.rfind("}") + 1])
        assert data["verdict"] in ("PASS", "FAIL")
        return data
    except Exception as e:
        return {"verdict": "ERROR", "rationale": f"judge error: {e}"}  # count errors separately
```

```python
# calibrate.py: agreement with human labels
from sklearn.metrics import cohen_kappa_score, confusion_matrix
human = [...]  # "PASS"/"FAIL" from experts
model = [...]  # judge verdicts on the same items (exclude ERROR, report them separately)
tn, fp, fn, tp = confusion_matrix(human, model, labels=["FAIL", "PASS"]).ravel()
print("TPR (passes caught):", tp / (tp + fn), " TNR (fails caught):", tn / (tn + fp))
print("Cohen's kappa:", cohen_kappa_score(human, model))
```

Many providers also support structured outputs or JSON modes that guarantee parseable responses; use them where available.

## Pairwise comparison done right

```text
Compare Answer A and Answer B for the question below on CONCISENESS WITHOUT LOSING
REQUIRED INFORMATION. Required points: {points}. Output "A", "B" or "TIE" with one sentence why.
```

Run each pair twice with A and B swapped. Count a win only if both runs agree; otherwise record a tie. Report win rate with a confidence interval.

## Worked example

A UK insurance comparison site used a 1–10 "helpfulness" judge; scores hovered around 8 regardless of prompt changes. Switching to four binary criteria (correct cover explanation, mentions exclusions, no advice beyond regulated scope, plain English) calibrated on 150 expert labels exposed that 12% of answers strayed into personal recommendations, a compliance concern the 1–10 score had hidden.

## Pitfalls

- **Vague criteria** ("is it good?").
- **Uncalibrated judges** reported as truth.
- **One judge for many criteria** in one prompt; split criteria for clarity.
- **Ignoring judge errors and cost**; log them.

## How to measure success

Each judge has a documented calibration (TPR, TNR, kappa) on a held-out human-labeled set, and is recalibrated when changed.

## Video lecture: LLM-as-judge: design, bias and calibration

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. LLM-as-judge
2. Analogy: exam markers
3. Judge formats
4. Prefer binary
5. Known biases
6. Mitigations
7. Calibrate the judge
8. Refine and repeat
9. Case: UK insurance comparison
10. Example: swap the order (illustrative)
11. Scenario: calibrating a service-charge judge (illustrative)
12. Common mistakes
13. Deeper: consistent-only arithmetic (illustrative)
14. Watch me do it: calibrating a judge (illustrative)
15. Recap

## Lecture transcript

### LLM-as-judge

Some qualities cannot be checked in code. Is this answer correct according to policy? Does the tone match the brand? Is the explanation actually helpful? For these, teams use a language model as a judge. Done well, it scales human judgment across thousands of outputs. Done badly, it produces confident numbers that mean nothing. In this lecture you will learn judge formats, known biases, and how to calibrate a judge against human experts.

### Analogy: exam markers

An analogy for judges. Think of a panel of exam markers grading essays. Give them a vague instruction like mark for quality, and one gives sevens, another gives nines, and the same marker changes their mind after lunch. Give them a clear rubric with examples of passing and failing answers, and check a sample of their marking against the head examiner, and marking becomes consistent and trustworthy. An LLM judge is a tireless marker. It still needs a rubric and a head examiner.

### Judge formats

There are four main formats. Pointwise with a rubric grades one output against explicit criteria. Reference-guided gives the judge a reference answer or key points to check against. Pairwise shows two outputs and asks which is better on a specific criterion, which is great when comparing an old and new prompt. And critique then verdict asks for a short rationale before the final decision, which tends to improve consistency and gives you readable reasons.

### Prefer binary

One strong recommendation: prefer binary pass or fail per criterion over one to ten scales. Numeric scales feel precise, but judges and humans use them inconsistently, and the difference between a six and a seven rarely means anything actionable. Binary criteria with explicit fail conditions are easier to calibrate, easier to explain and easier to act on.

### Known biases

Judges have known biases, documented in research like the MT-Bench study on judging LLM-as-a-judge, and in practice. Position bias: in pairwise comparisons, preferring whichever answer comes first, or second. Verbosity bias: preferring longer answers. Self-preference: favoring outputs from their own model family. Leniency: passing borderline cases. And criteria drift: reading vague criteria differently each time.

### Mitigations

Each bias has a mitigation. For position bias, run pairwise comparisons in both orders and count a win only when both agree. For verbosity, write criteria that reward concision. For self-preference, use a judge from a different model family where you can, and calibrate. For leniency and drift, use precise binary criteria with worked examples of pass and fail right in the prompt.

### Calibrate the judge

Now the most important idea: a judge is a model you are deploying, so evaluate it. Have domain experts label one to two hundred outputs as pass or fail. Run the judge on the same outputs. Measure agreement, but not just accuracy. Look at the true positive rate and the true negative rate separately, because a judge that passes everything looks accurate when most outputs are good. Add a chance-corrected statistic like Cohen's kappa.

### Refine and repeat

Then inspect the disagreements. Where the judge and experts differ, the rubric is usually ambiguous or missing an example. Refine it, and re-measure on a held-out split so you are not just tuning to the calibration set. And recalibrate whenever you change the judge model or prompt. The lesson gives you working Python for a binary rubric judge with rationale, error handling, and a calibration script.

### Case: UK insurance comparison

A real-world shaped example. A UK insurance comparison site used a one to ten helpfulness judge. Scores sat around eight no matter what they changed. They switched to four binary criteria: correct explanation of cover, mentions exclusions, no advice beyond their regulated scope, and plain English, calibrated on one hundred and fifty expert labels. That exposed a meaningful share of answers drifting into personal recommendations, a compliance issue the single score had completely hidden.

### Example: swap the order (illustrative)

A simple example of fixing position bias. You compare an old and a new prompt using a pairwise judge. Run one: new prompt shown first, it wins sixty percent. Run two, same pairs, swapped order: the new prompt wins only forty-five percent. The judge liked whatever came first. Now count only pairs where both runs agree. Maybe the new prompt wins thirty percent, loses twenty-five, and the rest are ties. That is a much more honest picture, and it might change your decision.

### Scenario: calibrating a service-charge judge (illustrative)

And a realistic calibration story, with illustrative numbers. A property portal in Dubai builds a judge for the criterion: does the answer state service charges only from the listing data? Experts label one hundred and fifty answers. The first judge version agrees with experts on most passing answers but catches only about half of the failing ones. Reading the disagreements shows the judge accepts ranges the model invented, like roughly twelve to fifteen dirhams per square foot. They add two fail examples with invented ranges to the rubric, and the judge's catch rate on failures rises sharply on a held-out set.

### Common mistakes

Common mistakes with judges. Asking one judge prompt to score five criteria at once. Using a one-to-ten scale and treating a shift from seven point two to seven point four as progress. Never checking the judge against humans. And silently counting judge errors, like unparseable output, as passes. Here is a quick question: if your judge started passing everything tomorrow, how would you find out?

### Deeper: consistent-only arithmetic (illustrative)

Let's deepen the pairwise example with the arithmetic. Suppose one hundred pairs. In the first ordering, the new prompt wins sixty. In the swapped ordering, forty-five. Count only pairs where both runs agree: say thirty wins, twenty-five losses, and forty-five ties. The net win rate over decisive pairs is thirty out of fifty-five. That still favors the new prompt, but far less than sixty percent suggested.

### Watch me do it: calibrating a judge (illustrative)

Watch me do it. I'm calibrating the policy-correctness judge from the lesson. First, the rubric: one criterion only, policy correctness, with an explicit fail condition, applying another country's policy, and two examples, a UAE seven-day pass and a UAE fourteen-day fail. The judge returns a rationale and a verdict as JSON, and errors are recorded separately. Second, I load one hundred and twenty answers that our support lead labeled pass or fail. Ninety are passes and thirty are fails. Third, I run the judge on all of them. Two come back as errors because the output was not valid JSON, so I set those aside and report them. Fourth, I run the calibration script. The true positive rate, passes correctly passed, is high, around ninety-six percent. The true negative rate, fails correctly caught, is only about sixty-five percent. Kappa is moderate. Fifth, I read the disagreements. Most missed fails are answers that are correct for the UK but given to a Saudi customer. The rubric examples only covered the UAE. I add one Saudi fail example and one line: check the country field before judging. Sixth, I rerun on a held-out batch the judge has never seen. The fail catch rate rises substantially. Numbers illustrative, but the process is exactly this.

### Recap

Recap. Use judges where code cannot decide. Prefer binary criteria with explicit examples, one criterion per judge prompt. Know the biases and mitigate them, especially by swapping order in pairwise tests. And always calibrate against human experts, reporting true positive and true negative rates and kappa. Your next step: pick one criterion, collect a hundred expert labels, and calibrate a judge using the code in the lesson.

## Key takeaways

- Prefer binary pass/fail per criterion, one criterion per judge prompt, with pass and fail examples.
- Mitigate position, verbosity, self-preference and leniency biases; swap order in pairwise comparisons.
- Calibrate judges against 100–200 expert labels, reporting TPR, TNR and Cohen's kappa.
- Recalibrate whenever the judge model or prompt changes; log judge errors separately.

## Try it

Collect 100 expert pass/fail labels for one criterion, run the judge from this lesson, and report TPR, TNR and kappa with three rubric improvements.

- [Previous: Code-based metrics and graders](https://optimizeall.com/learn/llm-evals-and-observability/code-based-metrics)
- [Next: Groundedness, safety, privacy and fairness metrics](https://optimizeall.com/learn/llm-evals-and-observability/groundedness-safety-and-quality)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
