Fine-Tuning, Distillation and Custom ModelsTraining data: design, synthesis and safety · Lesson 3 of 16
Dataset design and curation: quality beats quantity
Video lecture
Dataset design and curation: quality beats quantity
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Dataset design and curation
Here is a story I have seen many times. A team fine-tunes a model on twelve thousand examples, gets disappointing results, and spends two weeks tweaking learning rates. Nothing helps. Then someone reads a hundred training examples and finds duplicates, contradictory labels and a system prompt that does not match production. In this lesson you will learn to design and curate datasets so that the model learns what you actually mean.
0:31 Analogy: a recipe book
Think of training data as a recipe book you hand to a new cook. If half the recipes are duplicates, the cook thinks those dishes are twice as important. If two recipes for the same dish disagree on the salt, the cook learns to guess. And if the book only has summer dishes, the cook panics in winter. The cook can only be as good as the book. That is why dataset design matters more than any training setting.
1:05 Step 1: the labeling guide
Start with a labeling guide, before you collect anything. The task in one sentence and the exact output. Every label with a definition, two good examples and one tricky counter-example. Priority rules for ambiguous cases, like money issues outrank delivery issues. And what to do with out-of-scope, abusive or unexpected-language inputs. Then test the guide. Two people label the same hundred examples independently, and you measure agreement. If humans cannot apply the rule consistently, the model cannot learn it.
1:39 Step 2: coverage
Step two is coverage, not volume. Think in dimensions. Every intent or output type. Every language and script your users write, for example English, Arabic, Urdu and Roman Urdu. Every channel, because WhatsApp messages look nothing like emails. A range of difficulty, including messages with two issues or sarcasm. Realistic lengths. And the edge cases where production fails. Build a coverage matrix, with classes as rows and languages or channels as columns, and fill empty cells on purpose.
2:13 How many?
How many examples do you need? It depends on the task and the base model, so treat any number as orientation to verify. Format and style changes often show up with a few hundred good examples. Classification across many classes and languages often needs thousands. Complex reasoning needs more, or a different method altogether. But a thousand clean, consistent, representative examples beat ten thousand messy ones almost every time.
2:43 Step 3: clean
Step three, clean. Deduplicate exact and near duplicates, because duplicates overweight some patterns and leak across your splits. Review a sample of labels, starting with examples where a strong model disagrees with the label, since that is where errors cluster. Remove personal data you do not need. And normalize the format: the same system prompt and the same output formatting you will use in production, in every single example.
3:13 Step 4: split
Step four, split without leakage. You need a training set, a validation set for tuning decisions, and a test set you touch once at the end. For conversations, never split randomly. Split by customer or by time, so all messages from one customer, or all tickets after a date, live in one split. Otherwise almost identical conversations sit in both training and test, and your scores look far better than reality.
3:44 Hands-on: prep_data.py
The lesson's script does the mechanical part. It validates every line of your JSON lines file, checks that the last message is the assistant's label and that the label is allowed, removes duplicate user messages, splits by customer, strips metadata before writing the training files, and prints a table of class counts per split. Read that table. Every class should appear in every split, and nothing important should be vanishingly rare.
4:15 Worked example: Saudi telecom (illustrative)
A worked example, with illustrative numbers. A Saudi telecom started with twelve thousand complaint messages. Forty percent were near duplicates from automated notifications, and two teams labeled billing differently. They wrote a priority rule, relabeled fifteen hundred disputed messages, removed duplicates and split by customer. The dataset shrank to six thousand eight hundred examples, and the fine-tuned model's test accuracy rose substantially. Less data, better model. They documented it all in a data card: sources, dates, languages, counts, guide version, agreement, cleaning and known gaps.
4:52 Simple example: florist shop
A simple example. A florist's online shop wants to classify customer messages into order, delivery, complaint and other. They pull three hundred messages. Reading them, they find forty identical auto-replies, which they remove. They find two staff members labeled late delivery complaints differently, so they add a rule: anything about lateness is delivery. And they notice almost no messages from Valentine's week, their busiest period, so they pull more from last February. Three fixes, one afternoon, a much better dataset.
5:27 Read your data
One more habit that pays off: read your data. Before every training run, open fifty random examples and read them, start to finish. You will find things no script catches: a label that means something different to the night shift, a template reply pasted into hundreds of messages, a language you did not know you had. Keep a short log of what you found and what you changed. That log becomes part of your data card, and it is often the most useful page in it.
6:04 FAQ: only a few hundred examples?
A question from teams with little data: can we start fine-tuning with only a few hundred examples? For narrow format or style tasks, sometimes yes, especially with LoRA on a capable base model. The key is that those few hundred are clean, consistent and representative. Start there, evaluate honestly, and let the error analysis tell you where more data is needed. Adding targeted examples for the specific failures usually helps more than doubling the dataset at random.
6:37 Try this now
Try this now. Export fifty real examples for one task, and read every one of them. Count duplicates, count examples you are unsure how to label, and note which languages and channels appear. Those three numbers will tell you more about your fine-tuning readiness than any model benchmark.
6:58 Watch me do it
Watch me do it. I export eight hundred support messages with labels from two teams. First I read fifty. I spot auto-generated delivery notifications, so I run a dedup on normalized text: one hundred and ten duplicates removed. Then I sort by label and read twenty billing examples: team A labels late refunds as billing, team B as refund. I write a priority rule, refunds outrank billing, and relabel the sixty affected messages. Next the coverage matrix: rows are intents, columns are English, Arabic and Roman Urdu. Roman Urdu account-access has only four examples, so I flag it for collection. Then I run the prep script: it validates every line, splits by customer, and prints class counts per split. One rare class is missing from the test split entirely, so I adjust the split seed. Finally I fill in the data card: source, dates, counts, the new priority rule and the known gap.
8:05 Recap
Recap. The model becomes its data. Write and test a labeling guide. Cover the real variety of your traffic. Clean duplicates, fix labels and match production format. Split by customer or time. And document a data card. Your next step: write a one-page guide for one task, have two people label fifty examples, measure agreement, and run the prep script.
The model becomes its data
A fine-tuned model is a compressed copy of the examples you show it, including their mistakes, inconsistencies and gaps. Hyperparameters matter far less than data. Most failed fine-tunes are data failures: inconsistent labels, missing edge cases, leakage between training and test sets, or examples that do not look like real traffic.
Step 1: define the behavior precisely
Write a labeling guide before collecting data:
- The task in one sentence and the exact output format.
- Every label or field, with a definition, two positive examples and one tricky negative.
- Priority rules for ambiguous cases ("money issues outrank delivery issues").
- What to do when the input is out of scope, abusive or in an unexpected language.
Test the guide: have two people label the same 100 examples independently and measure agreement (for classification, Cohen's kappa is a common statistic). Low agreement means the guide, not the annotators, needs work. The model cannot learn a rule humans cannot apply consistently.
Step 2: collect representative examples
Aim for coverage, not volume:
| Dimension | Why it matters | Example |
|---|---|---|
| Intents / output types | Every class needs enough examples | 18 support intents |
| Languages and scripts | Models learn per-language behavior | English, Arabic, Urdu, Roman Urdu |
| Channels | Style differs by channel | WhatsApp vs email vs web form |
| Difficulty | Easy examples alone do not teach judgment | Multi-issue and sarcastic messages |
| Length | Train on the lengths you will see | 5-word texts and 300-word emails |
| Edge cases | Where production fails | Empty messages, screenshots described in text, mixed intents |
Build a coverage matrix (rows: classes; columns: languages or channels) and fill the gaps deliberately.
How many examples? It depends on the task and base model. As rough orientation (verify on your task): format and style changes often show effects with a few hundred good examples; classification with many classes and languages often needs thousands; complex reasoning behaviors need more, or a different method (RFT). Quality and coverage matter more than raw count.
Step 3: clean
- Deduplicate exact and near-duplicates (normalize whitespace and case; use embeddings or MinHash for near-duplicates). Duplicates overweight some patterns and leak across splits.
- Fix labels: sample 5–10% for review; prioritize examples where a strong model disagrees with the label.
- Remove or redact PII you do not need (next lessons).
- Normalize format: identical system prompt and chat template across examples; the same output formatting you want in production.
Step 4: split without leakage
Create train, validation (for tuning decisions) and test (touched once, at the end) sets. Split by entity or time, not just randomly: keep all messages from the same customer, or all tickets after a date, in one split. Otherwise near-identical conversations appear in both train and test and inflate scores.
Hands-on: validate, deduplicate and split
# prep_data.py: validate chat-format JSONL, dedupe, and split by customer (pip install pandas)
import json, hashlib, random, re
import pandas as pd
ALLOWED = {"refund", "delivery", "card_blocked", "account_access", "fees", "other"}
rows = []
for n, line in enumerate(open("raw.jsonl", encoding="utf-8"), 1):
try:
ex = json.loads(line)
msgs = ex["messages"]
assert msgs[-1]["role"] == "assistant", "last message must be the target"
assert msgs[-1]["content"] in ALLOWED, f"unknown label {msgs[-1]['content']}"
user = next(m["content"] for m in msgs if m["role"] == "user")
key = hashlib.sha1(re.sub(r"\s+", " ", user.strip().lower()).encode()).hexdigest()
rows.append({"customer": ex["customer_id"], "key": key, "label": msgs[-1]["content"], "example": ex})
except Exception as e:
print(f"line {n}: skipped ({e})")
df = pd.DataFrame(rows).drop_duplicates("key")
customers = sorted(df.customer.unique()); random.Random(42).shuffle(customers)
cut1, cut2 = int(0.8 * len(customers)), int(0.9 * len(customers))
split = {c: ("train" if i < cut1 else "val" if i < cut2 else "test") for i, c in enumerate(customers)}
df["split"] = df.customer.map(split)
for name, part in df.groupby("split"):
with open(f"{name}.jsonl", "w", encoding="utf-8") as f:
for ex in part.example:
ex = {"messages": ex["messages"]} # drop metadata such as customer_id from training files
f.write(json.dumps(ex, ensure_ascii=False) + "\n")
print(df.groupby(["split", "label"]).size().unstack(fill_value=0)) # check class balance per splitReview the printed table: every class should appear in every split, and no class should be vanishingly rare in training.
Worked example: a Saudi telecom's complaint classifier
The first dataset had 12,000 messages but 40% were near-duplicates from auto-generated notifications, and "billing" was labeled inconsistently between two teams. After writing a priority rule, relabeling 1,500 disputed messages, removing duplicates and splitting by customer, the dataset shrank to 6,800 examples, and test accuracy of the fine-tuned model rose substantially compared with the first attempt (numbers illustrative). Less data, better model.
Document it: a data card
Record source, time range, languages, class counts, labeling guide version, agreement score, cleaning steps, PII handling and known gaps. Future you (and auditors) will need it.
Pitfalls
- Random splitting of conversational data (leakage).
- Training on "clean" examples when production traffic is messy.
- Changing the system prompt between training and production.
- Ignoring class imbalance: a rare but important class ("fraud") gets ignored.
How to measure success
Annotator agreement above your threshold, a filled coverage matrix, zero duplicates across splits, a documented data card, and a validation set that predicts test performance.
Key takeaways
- Most failed fine-tunes are data failures, not hyperparameter failures
- Write and test a labeling guide; measure annotator agreement
- Cover classes, languages, channels, difficulty, length and edge cases
- Deduplicate and split by entity or time to avoid leakage
- Keep system prompts and formats identical between training and production; document a data card
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Write a one-page labeling guide for one task, have two people label the same 50 examples, measure agreement, and run the prep script on your data.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.