Fine-Tuning, Distillation and Custom ModelsTraining data: design, synthesis and safety · Lesson 3 of 16

Dataset design and curation: quality beats quantity

Article · 16 min · 8 min lecture

Video lecture

Dataset design and curation: quality beats quantity

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Dataset design and curation

  • The model becomes its data
  • Guide, collect, clean, split
  • Document it

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The model becomes its data

A fine-tuned model is a compressed copy of the examples you show it, including their mistakes, inconsistencies and gaps. Hyperparameters matter far less than data. Most failed fine-tunes are data failures: inconsistent labels, missing edge cases, leakage between training and test sets, or examples that do not look like real traffic.

Step 1: define the behavior precisely

Write a labeling guide before collecting data:

  • The task in one sentence and the exact output format.
  • Every label or field, with a definition, two positive examples and one tricky negative.
  • Priority rules for ambiguous cases ("money issues outrank delivery issues").
  • What to do when the input is out of scope, abusive or in an unexpected language.

Test the guide: have two people label the same 100 examples independently and measure agreement (for classification, Cohen's kappa is a common statistic). Low agreement means the guide, not the annotators, needs work. The model cannot learn a rule humans cannot apply consistently.

Step 2: collect representative examples

Aim for coverage, not volume:

DimensionWhy it mattersExample
Intents / output typesEvery class needs enough examples18 support intents
Languages and scriptsModels learn per-language behaviorEnglish, Arabic, Urdu, Roman Urdu
ChannelsStyle differs by channelWhatsApp vs email vs web form
DifficultyEasy examples alone do not teach judgmentMulti-issue and sarcastic messages
LengthTrain on the lengths you will see5-word texts and 300-word emails
Edge casesWhere production failsEmpty messages, screenshots described in text, mixed intents

Build a coverage matrix (rows: classes; columns: languages or channels) and fill the gaps deliberately.

How many examples? It depends on the task and base model. As rough orientation (verify on your task): format and style changes often show effects with a few hundred good examples; classification with many classes and languages often needs thousands; complex reasoning behaviors need more, or a different method (RFT). Quality and coverage matter more than raw count.

Step 3: clean

  • Deduplicate exact and near-duplicates (normalize whitespace and case; use embeddings or MinHash for near-duplicates). Duplicates overweight some patterns and leak across splits.
  • Fix labels: sample 5–10% for review; prioritize examples where a strong model disagrees with the label.
  • Remove or redact PII you do not need (next lessons).
  • Normalize format: identical system prompt and chat template across examples; the same output formatting you want in production.

Step 4: split without leakage

Create train, validation (for tuning decisions) and test (touched once, at the end) sets. Split by entity or time, not just randomly: keep all messages from the same customer, or all tickets after a date, in one split. Otherwise near-identical conversations appear in both train and test and inflate scores.

Hands-on: validate, deduplicate and split

# prep_data.py: validate chat-format JSONL, dedupe, and split by customer (pip install pandas)
import json, hashlib, random, re
import pandas as pd

ALLOWED = {"refund", "delivery", "card_blocked", "account_access", "fees", "other"}
rows = []
for n, line in enumerate(open("raw.jsonl", encoding="utf-8"), 1):
    try:
        ex = json.loads(line)
        msgs = ex["messages"]
        assert msgs[-1]["role"] == "assistant", "last message must be the target"
        assert msgs[-1]["content"] in ALLOWED, f"unknown label {msgs[-1]['content']}"
        user = next(m["content"] for m in msgs if m["role"] == "user")
        key = hashlib.sha1(re.sub(r"\s+", " ", user.strip().lower()).encode()).hexdigest()
        rows.append({"customer": ex["customer_id"], "key": key, "label": msgs[-1]["content"], "example": ex})
    except Exception as e:
        print(f"line {n}: skipped ({e})")

df = pd.DataFrame(rows).drop_duplicates("key")
customers = sorted(df.customer.unique()); random.Random(42).shuffle(customers)
cut1, cut2 = int(0.8 * len(customers)), int(0.9 * len(customers))
split = {c: ("train" if i < cut1 else "val" if i < cut2 else "test") for i, c in enumerate(customers)}
df["split"] = df.customer.map(split)

for name, part in df.groupby("split"):
    with open(f"{name}.jsonl", "w", encoding="utf-8") as f:
        for ex in part.example:
            ex = {"messages": ex["messages"]}          # drop metadata such as customer_id from training files
            f.write(json.dumps(ex, ensure_ascii=False) + "\n")
print(df.groupby(["split", "label"]).size().unstack(fill_value=0))   # check class balance per split

Review the printed table: every class should appear in every split, and no class should be vanishingly rare in training.

Worked example: a Saudi telecom's complaint classifier

The first dataset had 12,000 messages but 40% were near-duplicates from auto-generated notifications, and "billing" was labeled inconsistently between two teams. After writing a priority rule, relabeling 1,500 disputed messages, removing duplicates and splitting by customer, the dataset shrank to 6,800 examples, and test accuracy of the fine-tuned model rose substantially compared with the first attempt (numbers illustrative). Less data, better model.

Document it: a data card

Record source, time range, languages, class counts, labeling guide version, agreement score, cleaning steps, PII handling and known gaps. Future you (and auditors) will need it.

Pitfalls

  • Random splitting of conversational data (leakage).
  • Training on "clean" examples when production traffic is messy.
  • Changing the system prompt between training and production.
  • Ignoring class imbalance: a rare but important class ("fraud") gets ignored.

How to measure success

Annotator agreement above your threshold, a filled coverage matrix, zero duplicates across splits, a documented data card, and a validation set that predicts test performance.

Key takeaways

  • Most failed fine-tunes are data failures, not hyperparameter failures
  • Write and test a labeling guide; measure annotator agreement
  • Cover classes, languages, channels, difficulty, length and edge cases
  • Deduplicate and split by entity or time to avoid leakage
  • Keep system prompts and formats identical between training and production; document a data card

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Two annotators agree on only 60% of labels. What should you do first?
  2. Why split conversational data by customer or time rather than randomly?
  3. Your training examples use a different system prompt from production. What is the risk?

Put it into practice

Write a one-page labeling guide for one task, have two people label the same 50 examples, measure agreement, and run the prep script on your data.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.