Evaluating and Monitoring LLM ApplicationsMetrics and graders: code, judges and quality dimensions · Lesson 4 of 16

Code-based metrics and graders

Article · 13 min · 8 min lecture

Video lecture

Code-based metrics and graders

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Code-based metrics

  • Prefer code when code can decide
  • Exact, normalized, fuzzy
  • Limits of BLEU, ROUGE, similarity
  • A grader library

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Prefer code when code can decide

Before reaching for an LLM judge, ask: can a deterministic check decide this? Code-based graders are fast, free, reproducible and easy to debug. Many important properties are checkable in code:

PropertyCode-based check
Output is valid JSON matching a schemajsonschema / Pydantic validation
Classification label is correctExact match against expected label
Required fact presentRegex or normalized substring
Forbidden content absentRegex, denylist, PII detectors
Correct tool called with right argumentsCompare tool-call name and parsed args
Length and format constraintsWord or character counts, Markdown structure
SQL or code correctnessExecute against a test database or unit tests
Citation validityEvery cited ID exists in retrieved documents

Exact, normalized and fuzzy matching

  • Exact match works for labels, IDs and short structured answers. Normalize first: case, whitespace, punctuation, Unicode normalization (important for Arabic and Urdu text, where different code points can render identically).
  • Fuzzy match (edit distance, token overlap) tolerates minor variation, useful for names and addresses. Choose thresholds on real data.
  • Numeric tolerance: compare numbers with an absolute or relative tolerance, and parse currency formats carefully (1,000.50 vs 1.000,50).

Classic NLP metrics and their limits

BLEU and ROUGE measure n-gram overlap with a reference; embedding similarity (cosine similarity between vectors) measures semantic closeness. They are cheap, but for open-ended generation they correlate poorly with human judgments of quality: a correct answer phrased differently scores low, and a fluent wrong answer that reuses the reference's words scores high. Use them for narrow tasks (translation of fixed strings, near-duplicate detection) or as weak signals, not as release gates.

Key-point coverage: a robust middle ground

Instead of comparing to one reference answer, list the key points an answer must contain and forbidden points it must not. Some key points can be checked in code (a number, a URL, a policy name); others need a judge (Lesson 2). This handles many valid phrasings.

Hands-on: a small grader library

# graders.py: deterministic graders returning (score 0..1, reason)
import json, re, unicodedata
from jsonschema import validate, ValidationError

def normalize(s: str) -> str:
    s = unicodedata.normalize("NFKC", s).casefold()
    return re.sub(r"\s+", " ", s).strip()

def exact(output: str, expected: str):
    ok = normalize(output) == normalize(expected)
    return (1.0 if ok else 0.0, "exact match" if ok else f"expected {expected!r}")

def contains_all(output: str, required: list[str]):
    out = normalize(output)
    missing = [r for r in required if normalize(r) not in out]
    return (1 - len(missing) / max(len(required), 1), f"missing: {missing}" if missing else "all present")

def contains_none(output: str, forbidden_patterns: list[str]):
    hits = [p for p in forbidden_patterns if re.search(p, output, flags=re.I)]
    return (0.0 if hits else 1.0, f"forbidden: {hits}" if hits else "clean")

def valid_json(output: str, schema: dict):
    try:
        validate(json.loads(output), schema)
        return (1.0, "valid")
    except (json.JSONDecodeError, ValidationError) as e:
        return (0.0, f"invalid: {str(e)[:120]}")

def tool_called(tool_calls: list[dict], name: str, required_args: dict):
    for c in tool_calls:
        if c.get("name") == name and all(c.get("args", {}).get(k) == v for k, v in required_args.items()):
            return (1.0, f"{name} called correctly")
    return (0.0, f"{name} not called with {required_args}")

PHONE_PK = r"(\+92|0)3\d{2}[- ]?\d{7}"      # illustrative PII pattern: Pakistani mobile numbers
EMAIL = r"[\w.+-]+@[\w-]+\.[\w.]+"

Use it in tests:

from graders import contains_all, contains_none, EMAIL, PHONE_PK

def test_return_policy_answer():
    out = run_bot("Can I return shoes I bought 10 days ago? I'm in Lahore.")
    score, why = contains_all(out, ["14 days", "original packaging"])
    assert score == 1.0, why
    score, why = contains_none(out, [EMAIL, PHONE_PK])
    assert score == 1.0, why

Worked example: structured extraction

A logistics company in Karachi extracts shipment details (consignee, city, weight, COD amount) from WhatsApp messages. Their eval uses only code: schema validation, exact match on city (normalized list including Urdu and English spellings), numeric tolerance on weight, and exact match on COD amount in paisa. They reserved LLM judges for a small "notes" field. The eval runs in seconds on 500 items, so they run it on every prompt change.

Pitfalls

  • Brittle string checks that fail on harmless variation. Normalize, and test your grader on known-good and known-bad outputs.
  • Using similarity scores as pass/fail for open-ended answers.
  • Forgetting to grade tool calls, not just final text, for agents.
  • Graders with bugs. Unit-test your graders; a broken grader silently reports false quality.

How to measure success

Most checks that can be deterministic are deterministic, graders have their own unit tests, and the code-graded part of the suite runs in under a minute.

Key takeaways

  • Prefer deterministic checks whenever code can decide: schemas, labels, required and forbidden content, tool calls.
  • Normalize text (including Unicode for Arabic and Urdu) and use fuzzy or numeric tolerance where appropriate.
  • BLEU, ROUGE and similarity scores are weak signals for open-ended answers, not release gates.
  • Key-point coverage handles many valid phrasings; unit-test your graders.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which property is best checked with code rather than an LLM judge?
  2. A correct answer phrased differently from the reference gets a low ROUGE score. What does this show?

Put it into practice

Extend the grader library with two product-specific graders, unit-test them on known good and bad outputs, and replace any judge checks that code can decide.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.