---
title: "Synthetic data generation and distillation"
description: "Why synthetic data is now standard practice Real labeled data is slow and expensive to collect, and it rarely covers rare cases. Strong models can…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/synthetic-data-and-distillation
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Training data: design, synthesis and safety · lesson 4 of 16 · 16 min

# Synthetic data generation and distillation

## Why synthetic data is now standard practice

Real labeled data is slow and expensive to collect, and it rarely covers rare cases. Strong models can generate training data: new inputs, labels for unlabeled inputs, ideal responses, and preference pairs. **Distillation** uses a strong **teacher** model's outputs to train a smaller **student** model, keeping much of the teacher's quality on one task at a fraction of the cost and latency. Many public open-weight models were themselves trained partly on synthetic and distilled data.

Synthetic data is powerful and risky. It amplifies the teacher's biases and mistakes, can lack the messiness of real traffic, and has licensing and terms-of-service implications.

## Four ways to use a teacher

1. **Label real inputs.** You have real unlabeled messages; the teacher labels them; humans spot-check. Usually the safest and most valuable.
2. **Generate responses.** For real prompts, the teacher writes ideal answers in your required style (then humans review a sample).
3. **Generate inputs.** The teacher writes new realistic inputs to fill coverage gaps (rare intents, languages, edge cases).
4. **Generate preference pairs.** Sample several responses, have a judge (or humans) pick the better one, and create chosen/rejected pairs for DPO.

## Making synthetic inputs diverse

Asking "write 500 customer complaints" produces repetitive text. Instead, **control the attributes**: combine intent × language × channel × persona × difficulty × length, and generate a few examples per combination. Seed each prompt with one or two real examples (anonymized) to anchor style.

## Filtering is the real work

Generate more than you need and filter hard:

- **Validators:** schema checks, label in allowed set, length limits, language detection.
- **Deduplication:** exact and near-duplicate removal (synthetic data is especially repetitive).
- **Consistency checks:** have a second model (or the same model with a different prompt) re-label; discard disagreements or send them to humans.
- **Decontamination:** remove anything too similar to your evaluation set, or your scores will be inflated.
- **Human review:** sample every batch; track the rejection rate as a quality signal.

Keep a **real-data test set** that no synthetic example is derived from. Evaluate the student on real data only.

## License and terms checks (critical)

- **API terms of service.** Some model providers' terms restrict using outputs to develop models that compete with them. Read the current terms of any API you use as a teacher, and record your conclusion.
- **Open-weight teacher licenses.** Apache-2.0 and MIT teachers (for example gpt-oss, Qwen3/3.5 Apache releases, DeepSeek MIT releases) are simpler. The Llama licenses include conditions on using Llama outputs to train other models (for example naming requirements for distributed models); check the exact license version.
- **Input data rights.** Real inputs you label must be usable for training under your privacy notices and contracts (next lesson).

## Hands-on: generate, label and filter with a local teacher

This example uses an open-weight teacher served by vLLM (any OpenAI-compatible endpoint works).

```python
# synth.py: attribute-controlled generation + teacher labeling + filtering (pip install openai)
import itertools, json, os, random
from openai import OpenAI

teacher = OpenAI(base_url=os.getenv("TEACHER_URL", "http://localhost:8000/v1"), api_key=os.getenv("TEACHER_KEY", "none"))
MODEL = os.getenv("TEACHER_MODEL", "openai/gpt-oss-120b")
INTENTS = ["refund", "delivery", "card_blocked", "account_access", "fees", "other"]
LANGS = ["English", "Arabic", "Urdu", "Roman Urdu"]
CHANNELS = ["WhatsApp", "email"]
DIFFICULTY = ["simple", "two issues in one message", "angry and vague"]

def gen_input(intent, lang, channel, diff):
    prompt = (f"Write ONE realistic customer message to a digital bank.\nIntent: {intent}\nLanguage: {lang}\n"
              f"Channel: {channel}\nStyle: {diff}\nUse fictional names and numbers only. Output only the message.")
    r = teacher.chat.completions.create(model=MODEL, temperature=1.0, messages=[{"role": "user", "content": prompt}])
    return (r.choices[0].message.content or "").strip()

def label(text):
    r = teacher.chat.completions.create(model=MODEL, temperature=0, messages=[
        {"role": "system", "content": f"Classify into one of {INTENTS}. Money issues outrank delivery. Reply with the label only."},
        {"role": "user", "content": text}])
    return (r.choices[0].message.content or "").strip().lower()

seen, kept = set(), []
combos = list(itertools.product(INTENTS, LANGS, CHANNELS, DIFFICULTY)); random.shuffle(combos)
for intent, lang, channel, diff in combos:
    try:
        msg = gen_input(intent, lang, channel, diff)
        key = " ".join(msg.lower().split())
        if not (5 <= len(msg) <= 1200) or key in seen:
            continue
        seen.add(key)
        relabel = label(msg)
        if relabel != intent:            # disagreement: route to human review instead of training
            with open("review.jsonl", "a", encoding="utf-8") as f:
                f.write(json.dumps({"text": msg, "intended": intent, "teacher_label": relabel}, ensure_ascii=False) + "\n")
            continue
        kept.append({"messages": [{"role": "system", "content": "Classify the intent."},
                                  {"role": "user", "content": msg},
                                  {"role": "assistant", "content": intent}],
                     "meta": {"synthetic": True, "lang": lang, "channel": channel, "difficulty": diff}})
    except Exception as e:
        print("generation error:", e)

with open("synthetic.jsonl", "w", encoding="utf-8") as f:
    for ex in kept:
        f.write(json.dumps(ex, ensure_ascii=False) + "\n")
print(f"kept {len(kept)} of {len(combos)}; see review.jsonl for disagreements")
```

Then decontaminate against your test set (for example with embedding similarity), sample 10% for human review, and mix with real data. Track the synthetic share per class; keep real examples dominant where you have them.

## Worked example: distilling a support-reply style

A UK subscription box company used a strong API model with a long style guide for customer replies. They took 3,000 real (redacted) customer emails, generated teacher replies, had agents approve or edit 600 of them, filtered out replies failing a policy checker, and fine-tuned a small open model with LoRA on the approved and filtered set. On a held-out set of real emails graded by agents, the student matched the teacher's approval rate within a few points at a much lower cost per reply (illustrative). They kept the teacher as a fallback for low-confidence cases.

## Pitfalls

- Evaluating on synthetic data (circular and inflated).
- Uniform, too-clean synthetic inputs that do not match production.
- Ignoring teacher errors: a student faithfully learns them.
- Skipping the terms and license check for the teacher.

## How to measure success

A student evaluated on real held-out data meets your bar; the synthetic pipeline logs its filter and rejection rates; and the teacher's license/terms review is on file.

## Video lecture: Synthetic data generation and distillation

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Synthetic data and distillation
2. Analogy: master and apprentice
3. Distillation
4. Four uses
5. Diversity by design
6. Filter hard
7. License and terms
8. Hands-on: synth.py
9. Worked example: UK subscription box (illustrative)
10. Simple example: language school
11. Too much synthetic?
12. FAQ: will the student copy teacher mistakes?
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### Synthetic data and distillation

Imagine you need two thousand labeled customer messages in Roman Urdu for a rare complaint type, and you have forty. You could wait six months for real data. Or you could ask a strong model to help. In this lesson you will learn how to use synthetic data and distillation properly: the four ways to use a teacher model, how to make synthetic data diverse, how to filter it, and the license checks that keep you out of trouble.

### Analogy: master and apprentice

Here is an analogy for distillation. A master chef runs a famous restaurant, but hiring the master for every branch is too expensive. So the master cooks a thousand plates of the signature dish while an apprentice watches, tastes and practices. The apprentice will never be the master, but for that one dish, they can get very close, and they cost a fraction as much. The key word is one dish. Distillation shines on narrow tasks.

### Distillation

First, the idea of distillation. A strong teacher model produces outputs on your task, and a smaller student model is trained to imitate them. Done well, the student keeps much of the teacher's quality on that one task, at a fraction of the cost and latency. Many public open-weight models were themselves trained partly this way. But the student also inherits the teacher's mistakes and biases, so filtering is not optional.

### Four uses

There are four ways to use a teacher. One, label real inputs you already have, with human spot-checks. This is usually the safest and most valuable. Two, write ideal responses to real prompts in your required style. Three, generate new inputs to fill coverage gaps, like rare intents or languages. Four, generate preference pairs: several responses per prompt, a judge picks the better one, and you get chosen and rejected pairs for DPO.

### Diversity by design

If you simply ask for five hundred customer complaints, you get five hundred cousins of the same complaint. The fix is to control attributes. Combine intent, language, channel, persona, difficulty and length, and generate a few examples for each combination. For example, a refund request, in Roman Urdu, on WhatsApp, angry and vague. Seed prompts with one or two real, anonymized examples to anchor the style. Diversity comes from the grid, not from the temperature.

### Filter hard

Filtering is the real work. Validators check schema, allowed labels, length and language. Deduplication removes the repetition synthetic data loves. Consistency checks have the teacher, or a second model, re-label each message, and disagreements go to humans instead of training. Decontamination removes anything too similar to your evaluation set. And humans review a sample of every batch. Track the rejection rate; if it jumps, something upstream changed.

### License and terms

Now the license and terms check, and please do not skip it. Some API providers' terms restrict using outputs to develop models that compete with them. Read the current terms of any API you use as a teacher and write down your conclusion. Open-weight teachers under Apache or MIT, like gpt-oss, many Qwen releases or DeepSeek's MIT releases, are simpler. Llama licenses attach conditions to training other models on Llama outputs, such as naming requirements, so check the exact version. And your real inputs must be usable for training under your privacy notices and contracts.

### Hands-on: synth.py

The hands-on script uses an open-weight teacher behind any OpenAI-compatible endpoint. It loops over combinations of intent, language, channel and difficulty, generates one message each with fictional names and numbers, deduplicates, and has the teacher re-label each message at temperature zero with your priority rule. Agreements go into the training file with metadata marking them synthetic. Disagreements go to a review file for humans. Then you decontaminate, sample for review and mix with real data.

### Worked example: UK subscription box (illustrative)

A worked example, with illustrative numbers. A UK subscription box company used a big API model with a long style guide for customer replies. They generated teacher replies for three thousand real, redacted emails, had agents approve or edit six hundred, filtered out replies failing a policy checker, and trained a small open model with LoRA. On real held-out emails graded by agents, the student came within a few points of the teacher's approval rate at a much lower cost, with the teacher kept as a fallback.

### Simple example: language school

A simple example. A language school has two hundred real student questions about course schedules in English, but almost none in Arabic. They ask an open-weight teacher model to write Arabic versions of the questions in different styles, formal, casual and with typos, using made-up names. They generate six hundred, filter duplicates and anything off-topic, have a bilingual teacher review fifty, and keep four hundred. The Arabic gap in their coverage matrix is filled, and they test only on real Arabic questions collected later.

### Too much synthetic?

How much synthetic data is too much? There is no universal number, but watch two signals. First, per-class share: if a class is mostly synthetic, its test performance on real data is the only thing to trust. Second, style drift: synthetic messages tend to be cleaner, more polite and more grammatical than real ones. If your model starts struggling with typos, slang or voice-note transcripts, you have too much clean synthetic data. Keep real examples dominant wherever you have them.

### FAQ: will the student copy teacher mistakes?

A question I hear: if the teacher model makes mistakes, will the student copy them? Yes, faithfully. That is why filtering matters so much, and why the student should be evaluated against real data, not against the teacher. Sometimes a well-filtered student even outperforms its teacher on the narrow task, because the filters removed the teacher's worst outputs. But you only know that from an honest evaluation on real examples.

### Try this now

Try this now. Find the emptiest cell in your coverage matrix, the class and language combination with the fewest real examples. Write a generation prompt for it that specifies intent, language, channel and difficulty, and includes one anonymized real example. Generate twenty, and read all twenty before you generate any more.

### Watch me do it

Watch me do it. My coverage matrix shows almost no Roman Urdu messages about card blocking. I start the teacher model on our vLLM server and run the synthetic script only for that cell, across two channels and three difficulty styles, generating sixty messages with fictional names. The script deduplicates and re-labels each one with the teacher. Fifty-one agree with the intended label and go into the synthetic file; nine go to the review file. I read the nine: four are genuinely ambiguous, card blocked and fraud together, so I add them as edge cases with a human label. Five are off-topic and I drop them. Then I read twenty of the kept messages: two sound too polished for WhatsApp, so I add a style note to the generation prompt, typos and short sentences, and regenerate that batch. Finally I decontaminate against the test set and record the synthetic share for that class.

### Recap

Recap. Use teachers to label, write, fill gaps and create preference pairs. Design for diversity with attribute grids. Filter hard, and evaluate the student only on real held-out data. And always check terms and licenses. Your next step: run the pipeline for one under-covered cell of your coverage matrix, review fifty outputs, and report your rejection rate and the reasons.

## Key takeaways

- Use teachers to label real inputs, write responses, fill coverage gaps and create preference pairs
- Control attributes (intent × language × channel × difficulty) for diversity
- Filter hard: validators, dedup, re-label consistency, decontamination, human review
- Evaluate the student on real held-out data only
- Check API terms and teacher licenses before using outputs for training

## Try it

Run the synthetic pipeline for one under-covered class and language in your coverage matrix, review 50 outputs, and report your rejection rate and reasons.

- [Previous: Dataset design and curation: quality beats quantity](https://optimizeall.com/learn/fine-tuning-and-custom-models/dataset-design-and-curation)
- [Next: Training data safety, licensing and privacy](https://optimizeall.com/learn/fine-tuning-and-custom-models/data-safety-licensing-and-privacy)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
