Fine-Tuning, Distillation and Custom ModelsTraining data: design, synthesis and safety · Lesson 4 of 16

Synthetic data generation and distillation

Article · 16 min · 9 min lecture

Video lecture

Synthetic data generation and distillation

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Synthetic data and distillation

  • Four uses of a teacher
  • Diversity and filtering
  • License and terms checks

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why synthetic data is now standard practice

Real labeled data is slow and expensive to collect, and it rarely covers rare cases. Strong models can generate training data: new inputs, labels for unlabeled inputs, ideal responses, and preference pairs. Distillation uses a strong teacher model's outputs to train a smaller student model, keeping much of the teacher's quality on one task at a fraction of the cost and latency. Many public open-weight models were themselves trained partly on synthetic and distilled data.

Synthetic data is powerful and risky. It amplifies the teacher's biases and mistakes, can lack the messiness of real traffic, and has licensing and terms-of-service implications.

Four ways to use a teacher

  1. Label real inputs. You have real unlabeled messages; the teacher labels them; humans spot-check. Usually the safest and most valuable.
  2. Generate responses. For real prompts, the teacher writes ideal answers in your required style (then humans review a sample).
  3. Generate inputs. The teacher writes new realistic inputs to fill coverage gaps (rare intents, languages, edge cases).
  4. Generate preference pairs. Sample several responses, have a judge (or humans) pick the better one, and create chosen/rejected pairs for DPO.

Making synthetic inputs diverse

Asking "write 500 customer complaints" produces repetitive text. Instead, control the attributes: combine intent × language × channel × persona × difficulty × length, and generate a few examples per combination. Seed each prompt with one or two real examples (anonymized) to anchor style.

Filtering is the real work

Generate more than you need and filter hard:

  • Validators: schema checks, label in allowed set, length limits, language detection.
  • Deduplication: exact and near-duplicate removal (synthetic data is especially repetitive).
  • Consistency checks: have a second model (or the same model with a different prompt) re-label; discard disagreements or send them to humans.
  • Decontamination: remove anything too similar to your evaluation set, or your scores will be inflated.
  • Human review: sample every batch; track the rejection rate as a quality signal.

Keep a real-data test set that no synthetic example is derived from. Evaluate the student on real data only.

License and terms checks (critical)

  • API terms of service. Some model providers' terms restrict using outputs to develop models that compete with them. Read the current terms of any API you use as a teacher, and record your conclusion.
  • Open-weight teacher licenses. Apache-2.0 and MIT teachers (for example gpt-oss, Qwen3/3.5 Apache releases, DeepSeek MIT releases) are simpler. The Llama licenses include conditions on using Llama outputs to train other models (for example naming requirements for distributed models); check the exact license version.
  • Input data rights. Real inputs you label must be usable for training under your privacy notices and contracts (next lesson).

Hands-on: generate, label and filter with a local teacher

This example uses an open-weight teacher served by vLLM (any OpenAI-compatible endpoint works).

# synth.py: attribute-controlled generation + teacher labeling + filtering (pip install openai)
import itertools, json, os, random
from openai import OpenAI

teacher = OpenAI(base_url=os.getenv("TEACHER_URL", "http://localhost:8000/v1"), api_key=os.getenv("TEACHER_KEY", "none"))
MODEL = os.getenv("TEACHER_MODEL", "openai/gpt-oss-120b")
INTENTS = ["refund", "delivery", "card_blocked", "account_access", "fees", "other"]
LANGS = ["English", "Arabic", "Urdu", "Roman Urdu"]
CHANNELS = ["WhatsApp", "email"]
DIFFICULTY = ["simple", "two issues in one message", "angry and vague"]

def gen_input(intent, lang, channel, diff):
    prompt = (f"Write ONE realistic customer message to a digital bank.\nIntent: {intent}\nLanguage: {lang}\n"
              f"Channel: {channel}\nStyle: {diff}\nUse fictional names and numbers only. Output only the message.")
    r = teacher.chat.completions.create(model=MODEL, temperature=1.0, messages=[{"role": "user", "content": prompt}])
    return (r.choices[0].message.content or "").strip()

def label(text):
    r = teacher.chat.completions.create(model=MODEL, temperature=0, messages=[
        {"role": "system", "content": f"Classify into one of {INTENTS}. Money issues outrank delivery. Reply with the label only."},
        {"role": "user", "content": text}])
    return (r.choices[0].message.content or "").strip().lower()

seen, kept = set(), []
combos = list(itertools.product(INTENTS, LANGS, CHANNELS, DIFFICULTY)); random.shuffle(combos)
for intent, lang, channel, diff in combos:
    try:
        msg = gen_input(intent, lang, channel, diff)
        key = " ".join(msg.lower().split())
        if not (5 <= len(msg) <= 1200) or key in seen:
            continue
        seen.add(key)
        relabel = label(msg)
        if relabel != intent:            # disagreement: route to human review instead of training
            with open("review.jsonl", "a", encoding="utf-8") as f:
                f.write(json.dumps({"text": msg, "intended": intent, "teacher_label": relabel}, ensure_ascii=False) + "\n")
            continue
        kept.append({"messages": [{"role": "system", "content": "Classify the intent."},
                                  {"role": "user", "content": msg},
                                  {"role": "assistant", "content": intent}],
                     "meta": {"synthetic": True, "lang": lang, "channel": channel, "difficulty": diff}})
    except Exception as e:
        print("generation error:", e)

with open("synthetic.jsonl", "w", encoding="utf-8") as f:
    for ex in kept:
        f.write(json.dumps(ex, ensure_ascii=False) + "\n")
print(f"kept {len(kept)} of {len(combos)}; see review.jsonl for disagreements")

Then decontaminate against your test set (for example with embedding similarity), sample 10% for human review, and mix with real data. Track the synthetic share per class; keep real examples dominant where you have them.

Worked example: distilling a support-reply style

A UK subscription box company used a strong API model with a long style guide for customer replies. They took 3,000 real (redacted) customer emails, generated teacher replies, had agents approve or edit 600 of them, filtered out replies failing a policy checker, and fine-tuned a small open model with LoRA on the approved and filtered set. On a held-out set of real emails graded by agents, the student matched the teacher's approval rate within a few points at a much lower cost per reply (illustrative). They kept the teacher as a fallback for low-confidence cases.

Pitfalls

  • Evaluating on synthetic data (circular and inflated).
  • Uniform, too-clean synthetic inputs that do not match production.
  • Ignoring teacher errors: a student faithfully learns them.
  • Skipping the terms and license check for the teacher.

How to measure success

A student evaluated on real held-out data meets your bar; the synthetic pipeline logs its filter and rejection rates; and the teacher's license/terms review is on file.

Key takeaways

  • Use teachers to label real inputs, write responses, fill coverage gaps and create preference pairs
  • Control attributes (intent × language × channel × difficulty) for diversity
  • Filter hard: validators, dedup, re-label consistency, decontamination, human review
  • Evaluate the student on real held-out data only
  • Check API terms and teacher licenses before using outputs for training

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which synthetic-data use is usually safest and most valuable?
  2. Your student model scores 97% on a synthetic test set but 81% on real messages. What went wrong?
  3. Before using an API model as a teacher, what must you check?

Put it into practice

Run the synthetic pipeline for one under-covered class and language in your coverage matrix, review 50 outputs, and report your rejection rate and reasons.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.