---
title: "Code-based metrics and graders | Optimize All Academy"
description: "Prefer code when code can decide Before reaching for an LLM judge, ask: can a deterministic check decide this? Code-based graders are fast, free…"
url: https://optimizeall.com/learn/llm-evals-and-observability/code-based-metrics
updated: 2026-10-05
---

Evaluating and Monitoring LLM Applications · Metrics and graders: code, judges and quality dimensions · lesson 4 of 16 · 13 min

# Code-based metrics and graders

## Prefer code when code can decide

Before reaching for an LLM judge, ask: **can a deterministic check decide this?** Code-based graders are fast, free, reproducible and easy to debug. Many important properties are checkable in code:

| Property | Code-based check |
|---|---|
| Output is valid JSON matching a schema | `jsonschema` / Pydantic validation |
| Classification label is correct | Exact match against expected label |
| Required fact present | Regex or normalized substring |
| Forbidden content absent | Regex, denylist, PII detectors |
| Correct tool called with right arguments | Compare tool-call name and parsed args |
| Length and format constraints | Word or character counts, Markdown structure |
| SQL or code correctness | Execute against a test database or unit tests |
| Citation validity | Every cited ID exists in retrieved documents |

## Exact, normalized and fuzzy matching

- **Exact match** works for labels, IDs and short structured answers. Normalize first: case, whitespace, punctuation, Unicode normalization (important for Arabic and Urdu text, where different code points can render identically).
- **Fuzzy match** (edit distance, token overlap) tolerates minor variation, useful for names and addresses. Choose thresholds on real data.
- **Numeric tolerance**: compare numbers with an absolute or relative tolerance, and parse currency formats carefully (1,000.50 vs 1.000,50).

## Classic NLP metrics and their limits

**BLEU** and **ROUGE** measure n-gram overlap with a reference; **embedding similarity** (cosine similarity between vectors) measures semantic closeness. They are cheap, but for open-ended generation they correlate poorly with human judgments of quality: a correct answer phrased differently scores low, and a fluent wrong answer that reuses the reference's words scores high. Use them for narrow tasks (translation of fixed strings, near-duplicate detection) or as weak signals, not as release gates.

## Key-point coverage: a robust middle ground

Instead of comparing to one reference answer, list the **key points** an answer must contain and **forbidden points** it must not. Some key points can be checked in code (a number, a URL, a policy name); others need a judge (Lesson 2). This handles many valid phrasings.

## Hands-on: a small grader library

```python
# graders.py: deterministic graders returning (score 0..1, reason)
import json, re, unicodedata
from jsonschema import validate, ValidationError

def normalize(s: str) -> str:
    s = unicodedata.normalize("NFKC", s).casefold()
    return re.sub(r"\s+", " ", s).strip()

def exact(output: str, expected: str):
    ok = normalize(output) == normalize(expected)
    return (1.0 if ok else 0.0, "exact match" if ok else f"expected {expected!r}")

def contains_all(output: str, required: list[str]):
    out = normalize(output)
    missing = [r for r in required if normalize(r) not in out]
    return (1 - len(missing) / max(len(required), 1), f"missing: {missing}" if missing else "all present")

def contains_none(output: str, forbidden_patterns: list[str]):
    hits = [p for p in forbidden_patterns if re.search(p, output, flags=re.I)]
    return (0.0 if hits else 1.0, f"forbidden: {hits}" if hits else "clean")

def valid_json(output: str, schema: dict):
    try:
        validate(json.loads(output), schema)
        return (1.0, "valid")
    except (json.JSONDecodeError, ValidationError) as e:
        return (0.0, f"invalid: {str(e)[:120]}")

def tool_called(tool_calls: list[dict], name: str, required_args: dict):
    for c in tool_calls:
        if c.get("name") == name and all(c.get("args", {}).get(k) == v for k, v in required_args.items()):
            return (1.0, f"{name} called correctly")
    return (0.0, f"{name} not called with {required_args}")

PHONE_PK = r"(\+92|0)3\d{2}[- ]?\d{7}"      # illustrative PII pattern: Pakistani mobile numbers
EMAIL = r"[\w.+-]+@[\w-]+\.[\w.]+"
```

Use it in tests:

```python
from graders import contains_all, contains_none, EMAIL, PHONE_PK

def test_return_policy_answer():
    out = run_bot("Can I return shoes I bought 10 days ago? I'm in Lahore.")
    score, why = contains_all(out, ["14 days", "original packaging"])
    assert score == 1.0, why
    score, why = contains_none(out, [EMAIL, PHONE_PK])
    assert score == 1.0, why
```

## Worked example: structured extraction

A logistics company in Karachi extracts shipment details (consignee, city, weight, COD amount) from WhatsApp messages. Their eval uses only code: schema validation, exact match on city (normalized list including Urdu and English spellings), numeric tolerance on weight, and exact match on COD amount in paisa. They reserved LLM judges for a small "notes" field. The eval runs in seconds on 500 items, so they run it on every prompt change.

## Pitfalls

- **Brittle string checks** that fail on harmless variation. Normalize, and test your grader on known-good and known-bad outputs.
- **Using similarity scores as pass/fail** for open-ended answers.
- **Forgetting to grade tool calls**, not just final text, for agents.
- **Graders with bugs.** Unit-test your graders; a broken grader silently reports false quality.

## How to measure success

Most checks that *can* be deterministic *are* deterministic, graders have their own unit tests, and the code-graded part of the suite runs in under a minute.

## Video lecture: Code-based metrics and graders

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Code-based metrics
2. Analogy: gauges and craftspeople
3. Code can decide
4. Match robustly
5. Classic metrics mislead
6. Key-point coverage
7. Hands-on: graders.py
8. Case: Karachi shipment extraction
9. Pitfalls
10. Example: the JSON grader
11. Mistake: graders stricter than the business
12. Deeper: tolerance per field
13. Watch me do it: graders on one case
14. Recap

## Lecture transcript

### Code-based metrics

There is a temptation, once you discover LLM-as-judge, to use a model to grade everything. Resist it. Many of the most important properties of your system can be checked by a few lines of code: faster, free, perfectly reproducible and easy to debug. In this lecture you will learn which properties code can decide, how to match text robustly, why classic metrics like BLEU mislead, and how to build a small grader library you can trust.

### Analogy: gauges and craftspeople

Here is an analogy. In a factory, you do not ask an expert craftsperson to check whether every bolt is exactly ten millimeters. A gauge does that instantly, for free, all day. You save the expert for the things a gauge cannot judge, like whether the finish looks right. Code-based graders are your gauges. LLM judges are your expert craftspeople. Use gauges for everything a gauge can measure, and your experts will have time for the hard calls.

### Code can decide

Ask one question before choosing any grader: can a deterministic check decide this? Valid JSON that matches a schema. The correct classification label. A required fact present. A forbidden pattern absent, like an email address or phone number. The right tool called with the right arguments. Length and format rules. SQL that returns the right rows when executed. Citations that point to documents actually retrieved. All of these are code.

### Match robustly

Matching text well takes care. Normalize before comparing: case, whitespace, punctuation, and Unicode normalization. That last one matters for Arabic and Urdu, where different code points can render identically on screen but fail a naive comparison. Use fuzzy matching, like edit distance, for names and addresses, with thresholds tuned on real data. And compare numbers with tolerance, parsing currency formats carefully.

### Classic metrics mislead

Now a warning about classic metrics. BLEU and ROUGE count overlapping word sequences with a reference answer. Embedding similarity measures closeness in meaning. They are cheap, but for open-ended answers they correlate poorly with human judgment. A correct answer phrased differently scores low. A fluent, wrong answer that reuses the reference's words scores high. Use them for narrow tasks or as weak signals, never as a release gate.

### Key-point coverage

A robust middle ground is key-point coverage. Instead of comparing with a single reference answer, list the key points a good answer must contain, and the forbidden points it must not. Some key points can be checked in code, like a number, a policy name or a link. Others need a judge. This approach accepts the many valid phrasings real answers have, while still being specific about what matters.

### Hands-on: graders.py

The lesson gives you a small grader library in Python. Each grader returns a score between zero and one, plus a reason. Exact match with normalization. Contains all required phrases. Contains none of the forbidden patterns, with example patterns for emails and Pakistani mobile numbers. Valid JSON against a schema. And tool called with the required arguments. Reasons matter: when a test fails, you want to know why in one glance.

### Case: Karachi shipment extraction

Here is a team that got this right. A logistics company in Karachi extracts shipment details from WhatsApp messages: consignee, city, weight and cash-on-delivery amount. Their eval is almost entirely code: schema validation, exact match on city against a normalized list of Urdu and English spellings, numeric tolerance on weight, and exact match on amount in paisa. Only a free-text notes field uses a judge. It runs on five hundred items in seconds, so they run it on every prompt change.

### Pitfalls

Four pitfalls. Brittle string checks that fail on harmless variation, so normalize and test graders on known good and known bad outputs. Similarity scores used as pass or fail for open-ended answers. Grading only the final text of an agent and forgetting its tool calls. And graders with bugs. Unit-test your graders, because a broken grader silently reports false quality, which is worse than no grader.

### Example: the JSON grader

A simple example of a grader catching a real bug. Your bot must return JSON with an order status field. The valid JSON grader runs on a hundred test cases and fails three. You look: in all three, the customer wrote in Arabic, and the model wrapped its JSON in a friendly sentence first. That is a real production bug, found in seconds, at zero cost, with a clear reason attached. A judge might have scored those answers as helpful and missed the broken integration entirely.

### Mistake: graders stricter than the business

And a common mistake: writing graders that are stricter than the business needs. A team required the exact phrase fourteen days, and failed answers that said two weeks. Their pass rate looked terrible, and they spent a week changing prompts to satisfy a grader, not a customer. The fix was to normalize both forms and test the grader against a small set of known good and known bad answers first. Ask yourself: have you ever tested your graders themselves?

### Deeper: tolerance per field

One level deeper on numeric tolerance. Your extraction returns a weight of two point five kilograms, and the expected value is two point five zero. String equality fails, numeric comparison passes. But for money, use exact comparison in minor units, because a tolerance on currency hides real rounding bugs. Tolerance is a choice per field, not a global setting.

### Watch me do it: graders on one case

Watch me do it. I'm using the grader library on one real test case. The input: can I return shoes I bought ten days ago, I'm in Lahore. First, normalize. The function applies Unicode normalization, folds case and collapses spaces, so fourteen days with a capital letter or an extra space still matches. Second, contains all. Required phrases: fourteen days, and original packaging. The bot's answer mentions fourteen days but says in their box. Score: one half, and the reason line says missing original packaging. Is that a real failure or a strict grader? I check the policy: the requirement is unused items in original packaging, so the bot genuinely left something out. Third, contains none. Forbidden patterns: an email address and a Pakistani mobile number pattern. The answer is clean, score one. Fourth, for the order-status route, valid JSON. I paste a response that starts with Sure, here is your order, before the JSON. The parser fails and the reason shows the exact position. Fifth, tool called. The expected call is lookup order with the order ID, and the tool call list shows the right name and argument, so it passes. Finally, I write two unit tests for the graders themselves: one known good answer that must pass, and one known bad answer that must fail. Now I trust the numbers.

### Recap

Recap. Prefer code when code can decide. Normalize carefully, use fuzzy and numeric tolerance where appropriate, treat classic overlap metrics as weak signals, and use key-point coverage for open-ended answers. Your next step: take the grader library from the lesson, add two graders specific to your product, write unit tests for them, and move every check that can be deterministic out of your judge prompts.

## Key takeaways

- Prefer deterministic checks whenever code can decide: schemas, labels, required and forbidden content, tool calls.
- Normalize text (including Unicode for Arabic and Urdu) and use fuzzy or numeric tolerance where appropriate.
- BLEU, ROUGE and similarity scores are weak signals for open-ended answers, not release gates.
- Key-point coverage handles many valid phrasings; unit-test your graders.

## Try it

Extend the grader library with two product-specific graders, unit-test them on known good and bad outputs, and replace any judge checks that code can decide.

- [Previous: Building eval datasets: golden, production and synthetic](https://optimizeall.com/learn/llm-evals-and-observability/building-eval-datasets)
- [Next: LLM-as-judge: design, bias and calibration](https://optimizeall.com/learn/llm-evals-and-observability/llm-as-judge)
- [All lessons of Evaluating and Monitoring LLM Applications](https://optimizeall.com/learn/llm-evals-and-observability)
