---
title: "Building evaluation sets that reflect reality"
description: "Why evaluation is the core skill Prompt engineering without evaluation is guesswork. Every change (a new example, a reordered section, a different model)…"
url: https://optimizeall.com/learn/advanced-prompt-engineering/building-eval-sets
updated: 2026-10-05
---

Advanced Prompt Engineering · Evaluation sets and grading · lesson 12 of 17 · 12 min

# Building evaluation sets that reflect reality

## Why evaluation is the core skill

Prompt engineering without evaluation is guesswork. Every change (a new example, a reordered section, a different model) can improve some cases and break others. Without a fixed set of test cases and a way to score them, you are judging on the three outputs you happened to look at, which is how teams ship regressions.

Provider guidance on prompt engineering starts in the same place: define success criteria and a way to test against them *before* optimising the prompt.

## Step 1: define success criteria

Good criteria are **specific, measurable and tied to the real goal**. Compare:

- Vague: "The summary should be good."
- Specific: "Captures all decisions and owners from the meeting; no invented action items; under 200 words; readable by someone who missed the meeting."

Most tasks need several criteria across dimensions: correctness, completeness, format, tone, safety and cost or latency. Decide which are pass/fail gates (no invented facts) and which are scored (clarity from 1 to 5).

## Step 2: collect cases

A good evaluation set is a representative sample of reality plus deliberate stress tests:

- **Typical cases** drawn from real inputs (anonymised if they contain personal data).
- **Edge cases**: very long, very short, multilingual, messy formatting, missing information.
- **Known failures**: every bug report becomes a test case.
- **Adversarial cases**: injection attempts, false premises, out-of-scope requests.

Size depends on the stakes and how you grade. For early development, 20 to 50 well-chosen cases reveal most big problems. For decisions between two similar prompts, you need more, because small differences on small sets are often noise.

## Step 3: write reference answers or grading notes

For each case, record what a good answer must contain or must avoid. You do not always need a full gold answer; grading notes like "must mention the 14-day return window; must not promise a refund for opened software" are often more useful and less brittle.

```json
{
  "id": "refund-017",
  "input": "I opened the software box but it won't install. Refund?",
  "must_include": ["troubleshooting offer", "14-day window"],
  "must_not": ["promise of refund for opened software"],
  "tags": ["refund", "edge-case"]
}
```

## Step 4: keep it separate and versioned

- **Never put eval cases in your prompt** as examples. You will measure memorisation, not quality.
- **Version the set.** When you add cases, record it, so you know whether a score change came from the prompt or the test.
- **Tag cases** by category so you can see where a change helped or hurt, not just the overall average.

## Worked example: a meeting-notes assistant

A product team's first eval set had 25 cases, all clean English transcripts. Scores were high. In production, users uploaded transcripts with crosstalk, code-switching between Urdu and English, and speaker labels missing. The team added 15 such cases, and the score dropped noticeably, revealing the real work to be done. The lesson: an eval set that does not look like your real traffic is a confidence machine, not a measurement tool.

## Reading results honestly

- Look at failures, not just scores. Categorise them: retrieval miss, instruction ignored, format error, hallucination.
- Beware small differences. If prompt B beats prompt A by one case out of 40, you have not learned much. Rerun, since outputs vary between runs, and consider more cases.
- Check that improvements do not trade off hidden dimensions such as length, cost or tone.

## Failure modes

- **Teaching to the test.** Tuning the prompt until it passes these exact cases while real performance stagnates. Keep a held-out set you look at rarely.
- **Stale sets.** Your product changes, users change, and the eval set does not.
- **Only happy paths.** No unanswerable, adversarial or messy inputs.

## Hands-on: a minimal eval harness in Python

```python
import json, os, csv, time, anthropic
client = anthropic.Anthropic()

def run_prompt(system: str, user: str) -> str:
    resp = client.messages.create(
        model=os.environ.get("CLAUDE_MODEL", "claude-opus-5"), max_tokens=1024,
        system=system, messages=[{"role": "user", "content": user}],
    )
    return "".join(b.text for b in resp.content if b.type == "text")

def grade(case, output):
    text = output.lower()
    missing = [m for m in case["must_include"] if m.lower() not in text]
    violations = [m for m in case["must_not"] if m.lower() in text]
    return {"pass": not missing and not violations, "missing": missing, "violations": violations}

def evaluate(system_path, cases_path, out_path):
    system = open(system_path, encoding="utf-8").read()
    cases = [json.loads(line) for line in open(cases_path, encoding="utf-8")]
    rows = []
    for case in cases:
        t0 = time.perf_counter()
        output = run_prompt(system, case["input"])
        g = grade(case, output)
        rows.append({"id": case["id"], "tags": "|".join(case["tags"]), "pass": g["pass"],
                     "missing": ";".join(g["missing"]), "violations": ";".join(g["violations"]),
                     "latency_s": round(time.perf_counter() - t0, 2), "output": output})
    with open(out_path, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=rows[0].keys()); writer.writeheader(); writer.writerows(rows)
    by_tag = {}
    for r in rows:
        for t in r["tags"].split("|"):
            by_tag.setdefault(t, []).append(r["pass"])
    for t, v in sorted(by_tag.items()):
        print(f"{t:20} {sum(v)}/{len(v)}")

evaluate("prompts/support.v3.md", "evals/support.jsonl", "results/support.v3.csv")
```

Keyword grading is crude (it misses paraphrases), so treat it as a first gate and add rubric grading from the next lesson for nuanced criteria.

## Using an eval framework

Open-source tools such as Promptfoo let you define prompts, providers and test cases in a config file and run comparisons across models from the command line; provider consoles (for example the Anthropic Console and the OpenAI platform) also offer evaluation features. A minimal Promptfoo config looks like this; check the current docs for provider IDs and assertion types:

```yaml
# promptfooconfig.yaml
prompts:
  - file://prompts/support.v3.md
providers:
  - anthropic:messages:claude-opus-5
tests:
  - vars: { message: "I opened the software box but it won't install. Refund?" }
    assert:
      - type: icontains
        value: "14 days"
      - type: llm-rubric
        value: "Does not promise a refund for opened software"
```

## Generating synthetic cases carefully

When real data is scarce, a model can draft extra cases ("write 20 messy customer messages about late deliveries, mixing English and Urdu"). Have a human review them, tag them as synthetic, and never let synthetic cases outnumber real ones in the set you use for decisions.

## Going further

Automate the loop: a script that runs every case through the current prompt, applies graders, and outputs a table by tag, plus a diff against the last run. Many open-source and commercial evaluation tools exist; the tool matters less than the discipline of running it on every change.

## Video lecture: Building evaluation sets that reflect reality

Lecture coming soon · 11 chapters · about 7 minutes. Read the full transcript below.

1. Building evaluation sets
2. Evaluation is the core skill
3. Success criteria
4. Collecting cases
5. Grading notes
6. Keep it honest
7. Running evals
8. Read results honestly
9. Example 1: product titles eval
10. Example 2: bank card assistant (illustrative)
11. Recap

## Lecture transcript

### Building evaluation sets

A product team's meeting-notes assistant scored beautifully on its test set. Then real users uploaded transcripts with crosstalk, missing speaker labels, and people switching between Urdu and English mid-sentence. Quality fell off a cliff. The test set had been measuring a world that did not exist. In this lecture you will learn to build evaluation sets that reflect reality: defining success criteria, collecting the right cases, writing grading notes, keeping sets honest, and running them with a small harness or an eval framework.

### Evaluation is the core skill

Prompt engineering without evaluation is guesswork. Every change, a new example, a reordered section, a different model, can improve some cases and break others. Without a fixed set of test cases and a way to score them, you are judging on the three outputs you happened to look at. That is exactly how teams ship regressions. Provider guidance starts in the same place: define success criteria and a way to test against them before optimising the prompt.

### Success criteria

Good criteria are specific, measurable and tied to the real goal. The summary should be good is vague. Captures all decisions and owners; no invented action items; under two hundred words; readable by someone who missed the meeting: that is specific. Most tasks need criteria across several dimensions: correctness, completeness, format, tone, safety, and cost or latency. Decide which are pass or fail gates, like no invented facts, and which are scored, like clarity from one to five.

### Collecting cases

Now collect cases. Start with typical cases drawn from real inputs, anonymised if they contain personal data. Add edge cases: very long, very short, multilingual, messy formatting, missing information. Turn every bug report into a test case. And add adversarial cases: injection attempts, false premises, out-of-scope requests. For early development, twenty to fifty well-chosen cases reveal most big problems. To decide between two similar prompts, you need more, because small differences on small sets are often just noise.

### Grading notes

For each case, record what a good answer must contain or must avoid. You do not always need a full gold answer. Grading notes are often more useful and less brittle. For a refund question about opened software, the notes might say: must include a troubleshooting offer and the fourteen-day window; must not promise a refund for opened software; tags refund and edge case. Tags matter, because they let you see where a change helped or hurt, not just the overall average.

### Keep it honest

Keep the set honest. Never put eval cases in your prompt as examples, or you will measure memorisation. Version the set, so you know whether a score change came from the prompt or the test. Keep a held-out set you look at rarely, to avoid teaching to the test. And update it as your product and users change, because stale sets measure yesterday. When real data is scarce, a model can draft synthetic cases, like twenty messy customer messages about late deliveries mixing English and Urdu. Have a human review them, tag them as synthetic, and never let them outnumber real cases.

### Running evals

The lesson's hands-on harness is about forty lines of Python. It reads your system prompt from a file and your cases from a JSON lines file, runs each case through the model, grades with must-include and must-not checks, records latency, writes a CSV of results, and prints pass rates by tag. Keyword grading is crude, it misses paraphrases, so treat it as a first gate, and add rubric grading from the next lesson. If you prefer a framework, open-source tools like Promptfoo let you define prompts, providers and tests with assertions in a config file, and provider consoles offer evaluation features too.

### Read results honestly

Read results honestly. Look at failures, not just scores, and categorise them: retrieval miss, instruction ignored, format error, hallucination. Beware small differences; if prompt B beats prompt A by one case out of forty, you have not learned much, so re-run, because outputs vary between runs. And check that improvements do not trade off hidden dimensions like length, cost or tone. Back to our meeting-notes team: adding fifteen realistic, messy cases dropped their score, and revealed the real work to be done.

### Example 1: product titles eval

A simple worked example. You have a prompt that writes product titles for an online shop. Your first eval is five products you picked, all neat and in English. Every title looks great. Now build it properly. Ten typical products from last month's real catalogue. Five edge cases: a product with a very long name, one with no brand, one in Arabic, one with a size range, one bundle. Three known failures from complaints. And two adversarial inputs, like a product name containing an instruction. Twenty cases, each with grading notes such as must include the size, must not exceed sixty characters. Suddenly you see two real problems the first set hid.

### Example 2: bank card assistant (illustrative)

Now a business scenario, with illustrative numbers. A bank's customer-service team in Riyadh is building an assistant for card questions. Their first eval had thirty questions written by the product team, and scored ninety-five percent. Before launch, they sampled one hundred and fifty real, anonymised customer messages from the last quarter, in Arabic, English and mixed, including typos and long complaints. They tagged each by topic and difficulty, wrote grading notes with compliance input, and added fifteen adversarial prompts. Scores on the new set: seventy-eight percent overall, but only sixty-one percent on mixed-language messages and on card-dispute questions. That told them exactly where to work. Three prompt iterations later they reached ninety percent on the realistic set, with no tag below eighty-two. Illustrative figures, but notice which set told the truth.

### Recap

To recap. Define specific, measurable criteria before you optimise. Build sets from real inputs plus edge, failure and adversarial cases, with grading notes and tags. Keep cases out of prompts, version the set, keep a held-out slice, and read failures, not just scores. Try this now: write five success criteria and twenty-five tagged test cases for one AI task you care about, and keep five of them as a held-out set. Next: grading outputs with code checks, rubrics and LLM judges.

## Key takeaways

- Define specific, measurable success criteria before optimising any prompt.
- Build eval sets from real inputs plus edge, failure and adversarial cases; tag them by category.
- Grading notes (must include / must not) are often more robust than full gold answers.
- Keep eval cases out of prompts, version the set, keep a held-out set, and read failures, not only scores.

## Try it

Write 5 success criteria and 25 tagged test cases (typical, edge, failure, adversarial) for one AI task you care about. Keep 5 as a held-out set.

- [Previous: Prompt injection and defence basics](https://optimizeall.com/learn/advanced-prompt-engineering/prompt-injection-defence)
- [Next: Grading outputs: code checks, rubrics and LLM-as-judge](https://optimizeall.com/learn/advanced-prompt-engineering/grading-rubrics-llm-judge)
- [All lessons of Advanced Prompt Engineering](https://optimizeall.com/learn/advanced-prompt-engineering)
