Skip to content

Evaluating and Monitoring LLM Applications · Foundations: success criteria, error analysis and datasets · lesson 3 of 16 · 14 min

Building eval datasets: golden, production and synthetic

Datasets are the backbone

An eval is only as good as its dataset. A brilliant judge scoring the wrong inputs tells you nothing. Build datasets deliberately, from three sources, and treat them as versioned, owned assets.

Three sources of examples

1. Golden datasets (hand-curated). Inputs with reference answers or labels written or verified by domain experts. Small (often 50–300 items per task to start), high quality, slow to build. Use for release gates and judge calibration.

2. Production samples. Real user inputs from logs, sampled and labeled. The best reflection of reality, including messy phrasing, typos, code-switching (for example Urdu-English or Arabic-English), and unexpected requests. Requires privacy handling.

3. Synthetic data. Inputs generated by an LLM to fill coverage gaps: rare intents, edge cases, adversarial inputs, languages you have few samples for. Fast and cheap, but prone to being too clean and repetitive.

Generating synthetic data well

Naive prompting ("generate 100 customer questions") produces bland, similar questions. Use dimensions to structure variety:

# synth_inputs.py: structured synthetic test inputs (review before use)
import itertools, json, os
from anthropic import Anthropic

client = Anthropic()  # reads ANTHROPIC_API_KEY from the environment
MODEL = os.environ["EVAL_GEN_MODEL"]  # set to a current model ID from the provider's docs

intents = ["return request", "delivery delay", "wrong size", "payment failed", "cancel order"]
personas = ["first-time buyer, polite", "repeat customer, frustrated", "non-native English speaker"]
countries = ["PK", "AE", "GB"]
complications = ["none", "order is outside return window", "asks about another person's order"]

rows = []
for intent, persona, country, comp in itertools.product(intents, personas, countries, complications):
    prompt = (f"Write one realistic customer support message for an online fashion store.\n"
              f"Intent: {intent}. Persona: {persona}. Customer country: {country}. Complication: {comp}.\n"
              f"Write only the customer's message, 1-3 sentences, natural and imperfect.")
    try:
        msg = client.messages.create(model=MODEL, max_tokens=200,
                                     messages=[{"role": "user", "content": prompt}])
        text = msg.content[0].text.strip()
    except Exception as e:  # network or API error: record and continue
        text, comp = f"GENERATION_FAILED: {e}", comp
    rows.append({"intent": intent, "persona": persona, "country": country,
                 "complication": comp, "input": text, "source": "synthetic"})

with open("synthetic_inputs.jsonl", "w", encoding="utf-8") as f:
    for r in rows:
        f.write(json.dumps(r, ensure_ascii=False) + "\n")
print(len(rows), "rows")

The dimension tags (intent, persona, country, complication) let you later slice results: "we fail on AE + outside-return-window". Always review synthetic samples and discard unrealistic ones. Generate inputs synthetically; be much more careful generating expected outputs, which should be verified by experts for golden sets.

Labels and references

Depending on the metric, each example may need:

  • A reference answer (for comparison-based grading).
  • Required facts or key points ("mentions 14-day window", "offers exchange").
  • Forbidden content ("must not reveal order details without verification").
  • Expected tool calls (for agents): name and key arguments.
  • Metadata for slicing: language, country, intent, difficulty, source.

Key-point lists are often more robust than full reference answers because many phrasings are acceptable.

Sizing and coverage

There is no magic number. Practical guidance:

  • Start small and useful (tens to low hundreds per critical task) rather than large and noisy.
  • Ensure coverage of your failure taxonomy and important slices (languages, regions, customer types).
  • For comparing two systems, larger sets reduce noise; check confidence intervals (Module 4) before trusting small differences.
  • Keep separate adversarial sets for must-never-happen failures.

Versioning and hygiene

  • Store datasets in version control (JSONL/CSV) or a platform with versioning; record which version produced which scores.
  • Avoid contamination: do not paste eval items into prompts as few-shot examples; keep a held-out set that is never used during prompt development to detect overfitting to your own eval.
  • Privacy: production samples may contain personal data. Minimize, redact or pseudonymize; restrict access; follow retention rules and applicable laws (for example UK GDPR, EU GDPR, UAE and Saudi PDPL). Document the lawful basis with your privacy team.
  • Refresh with new production failures monthly.

Worked example: a bilingual golden set

A Riyadh bank's assistant needed Arabic and English coverage. The team built 200 golden items: 120 from anonymized production chats (reviewed by compliance), 80 synthetic for rare intents (card blocking abroad, disputed transfers). Every item had key points and forbidden-content lists reviewed by a product specialist. A held-out set of 50 items was locked away and used only at release time.

Pitfalls

  • Synthetic-only datasets that miss real user phrasing.
  • Reference answers written by the model you are testing.
  • No slicing metadata, so you cannot see where failures cluster.
  • Stale datasets that no longer reflect product or policy.

How to measure success

Coverage of every top failure category and key slice, a documented version history, and a held-out set used only for release decisions.

Video lecture: Building eval datasets: golden, production and synthetic

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

  1. Building eval datasets
  2. Analogy: the driving-test route
  3. Two sources
  4. Synthetic data
  5. Hands-on: dimension-driven generation
  6. Labels per example
  7. Sizing and coverage
  8. Hygiene
  9. Case: bilingual bank golden set
  10. Example: one well-labeled item
  11. Common mistakes
  12. Deeper: how contamination sneaks in
  13. Watch me do it: dataset v1
  14. Recap

Lecture transcript

Building eval datasets

An eval is only as good as its dataset. A brilliant judge scoring the wrong inputs tells you nothing useful. In this lecture you will learn the three sources of eval examples, golden, production and synthetic, how to generate synthetic data that is actually diverse, what labels each example needs, and how to keep datasets versioned, private and uncontaminated.

Analogy: the driving-test route

An analogy for eval datasets. A driving test is only meaningful if the route includes roundabouts, hill starts, parallel parking and a busy junction. If the examiner only ever drove you around an empty parking lot, passing would mean nothing. Your eval dataset is the test route. Golden examples are the standard maneuvers, production samples are real traffic, and synthetic data adds the rare situations, like a sudden rainstorm, that you might not meet for months.

Two sources

Golden datasets are hand-curated by domain experts, with reference answers or labels you trust. They are small, often fifty to a few hundred items per task to start, high quality and slow to build. Use them for release gates and to calibrate judges. Production samples are real user inputs from your logs, complete with typos and code-switching, like Urdu mixed with English, or Arabic with English. They are the best mirror of reality, and they need careful privacy handling.

Synthetic data

The third source is synthetic data, generated by a model to fill coverage gaps: rare intents, edge cases, adversarial inputs and languages where you have few samples. It is fast and cheap, but naive prompting produces bland, repetitive questions. The fix is to structure variety with dimensions. For a support bot: intent, persona, customer country and a complication, like an order outside the return window, or someone asking about another person's order.

Hands-on: dimension-driven generation

The lesson includes a script that loops over those dimensions and asks a model for one realistic, imperfect customer message per combination. Each row keeps its dimension tags. That pays off later, when you can say precisely: we fail on Emirati customers whose orders are outside the return window. Always review synthetic samples and discard unrealistic ones. And be very careful generating expected answers synthetically. For golden sets, experts verify the answers.

Labels per example

What labels does each example need? It depends on the metric. A reference answer, for comparison grading. Required key points, like mentions the fourteen-day window. Forbidden content, like must not reveal order details without verification. Expected tool calls for agents. And metadata for slicing: language, country, intent, difficulty and source. Key-point lists are often more robust than a single reference answer, because many phrasings are acceptable.

Sizing and coverage

How big should a dataset be? There is no magic number. Start small and useful, tens to low hundreds per critical task. Make sure you cover every top category in your failure taxonomy and every important slice, like languages and regions. When comparing two systems, bigger sets reduce noise, so check confidence intervals before trusting small differences. And keep separate adversarial sets for failures that must never happen.

Hygiene

Now hygiene. Version datasets in git or a platform, and record which version produced which scores. Avoid contamination: never paste eval items into prompts as examples, and keep a held-out set that is only used at release time, to catch overfitting to your own eval. Handle production data lawfully: minimize, redact or pseudonymize, restrict access, and follow your privacy team's guidance under laws such as UK GDPR or the data protection laws in the UAE and Saudi Arabia. And refresh monthly with new failures.

Case: bilingual bank golden set

A worked example from a bank in Riyadh. Its assistant needed Arabic and English coverage. The team built two hundred golden items: one hundred and twenty from anonymized production chats reviewed by compliance, and eighty synthetic items for rare intents like blocking a card abroad. Every item had key points and forbidden content reviewed by a product specialist. Fifty more items were locked away and used only for release decisions.

Example: one well-labeled item

A simple example of a well-labeled item. Input: I bought shoes in Dubai eight days ago, can I return them? Country: UAE. Key points: seven-day return window, offer exchange or store credit if eligible. Forbidden: promising a refund. Expected tool call: none. Slice tags: returns, UAE, English, edge case past window. Source: golden, verified by the support lead. With that one row, a code check, a judge and a slice report can all do their jobs.

Common mistakes

Common mistakes with datasets. Building only synthetic data, which misses how real customers actually write. Letting the model under test write its own expected answers. Forgetting slicing metadata, so you cannot see that failures cluster in one country. And never refreshing, so the dataset slowly describes last year's product. Try this now: open your current eval set and count how many items came from real production traffic. If the answer is zero, that is your next task.

Deeper: how contamination sneaks in

One level deeper on contamination. A prompt engineer adds three great examples to the system prompt. They are word for word from the golden set. Scores jump five points overnight. Nothing about the product improved; the model was simply shown the answers. The rule: examples used in prompts must never come from the eval files, and the held-out set catches this kind of accident.

Watch me do it: dataset v1

Watch me do it. I run the synthetic input script with four dimensions: five intents, three personas, three countries and three complications. That is one hundred and thirty-five combinations, one message each. The model is set by an environment variable, and failed calls are recorded instead of crashing the run. When it finishes, I read a sample of thirty. Most are good, like a frustrated repeat customer in Lahore asking why a delivery is late. Four are unrealistic, for example a first-time buyer writing an oddly formal legal letter about a wrong size. I delete those and note that the persona prompt needs a length limit. Next I label the ones I keep that target my top failure category, wrong-country policy. For each I add key points, like the seven-day window for the UAE, and forbidden content, like any mention of the fourteen-day UK window. I tag the source as synthetic. Then I add twenty real production messages from the same category, redacted, with the support lead verifying expected key points. Now the file has fifty items for that category, slice tags on every row, and a mix of sources. I commit it as dataset version one, and I move ten items to a held-out file that nobody uses while tuning prompts.

Recap

Recap. Combine golden, production and synthetic examples. Structure synthetic variety with dimensions and review everything. Label with key points, forbidden content and metadata. Cover your taxonomy, version your data, avoid contamination and handle privacy properly. Your next step: build a first dataset of fifty items for your top failure category, with slicing metadata, and commit it to version control.

Key takeaways

  • Combine expert-curated golden sets, privacy-safe production samples and structured synthetic data.
  • Structure synthetic variety with dimensions (intent, persona, region, complication) and review outputs.
  • Label with key points, forbidden content, expected tool calls and slicing metadata.
  • Version datasets, keep a held-out set, avoid contamination and handle personal data lawfully.

Try it

Create a 50-item dataset for your top failure category with key points, forbidden content and slicing metadata, and commit it as v1.