Advanced Prompt EngineeringFew-shot examples and structured outputs · Lesson 3 of 17

Few-shot and example design

Article · 12 min · 8 min lecture

Video lecture

Few-shot and example design

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

Few-shot and example design

  • Why examples are powerful
  • Four properties of good examples
  • Boundary cases and dynamic selection
  • Measuring value

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why examples are so powerful

Instructions describe what you want; examples show it. For tone, format, edge-case handling and classification boundaries, a few well-chosen examples often do more than a page of rules. Provider guidance commonly suggests a small handful (roughly three to five) of diverse, relevant examples as a strong starting point.

The mental model: the model generalises from the pattern your examples share, including patterns you did not intend. If all your examples are about software companies, outputs will drift toward software framing. If all examples are three sentences long, you have just set a length rule.

The four properties of good examples

  1. Relevant. They look like your real inputs, including messiness: typos, mixed languages, long rambling messages.
  2. Diverse. They cover the main categories and at least one tricky boundary case. Vary length, topic and phrasing so the model does not latch onto surface features.
  3. Correct. A single wrong label in your examples teaches the wrong rule. Have a domain expert check them.
  4. Clearly separated. Wrap them in tags such as <example> so the model does not confuse examples with the actual input.

A worked example: classifying lead intent

<instructions>
Classify each inbound message as HOT, WARM, COLD or SPAM.
HOT = named budget or timeline within 90 days.
WARM = clear need, no budget or timeline.
COLD = general curiosity, students, job seekers.
Return only the label.
</instructions>

<examples>
<example>
<input>We're launching in Riyadh in March and need a paid social partner. Budget approved.</input>
<output>HOT</output>
</example>
<example>
<input>hi, do u do tiktok ads? thinking about it for next yr</input>
<output>WARM</output>
</example>
<example>
<input>I'm writing my dissertation on influencer marketing, could I interview you?</input>
<output>COLD</output>
</example>
<example>
<input>Congratulations! Claim your free SEO audit now at this link</input>
<output>SPAM</output>
</example>
</examples>

<input>{{message}}</input>

Notice the second example: informal, lowercase, vague timeline. It teaches the WARM boundary more effectively than another sentence in the definition.

Designing boundary examples

The most valuable examples sit on decision boundaries: the message that looks HOT but lacks a timeline, or the complaint that looks like spam but is a real customer. To find them:

  • Run your prompt on 50 to 100 real inputs.
  • Collect the cases where the model and a human disagree.
  • Promote the most instructive disagreements into examples, with the correct label.

This is a loop, not a one-time task. Your example set should evolve with your evaluation set (covered in the evaluation module), but never copy your test cases into the prompt, or you will fool yourself about accuracy.

Examples with reasoning

For judgement-heavy tasks, include a short rationale in each example:

<example>
<input>Our Dubai store needs this by Eid, can you quote?</input>
<reasoning>Implied timeline within 90 days and explicit request for a quote.</reasoning>
<output>HOT</output>
</example>

This teaches the way of deciding, not only the answer. With reasoning-capable models, showing the reasoning style in examples tends to carry over to how the model thinks about new cases. If your pipeline needs just the label, ask for the reasoning in a separate tag and parse out the output.

Failure modes

  • Overfitting to surface features. All positive examples mention a currency, so the model treats any currency mention as HOT. Fix with contrasting examples.
  • Label imbalance. Three HOT examples and one COLD can bias toward HOT. Roughly balance, or match your real distribution deliberately.
  • Example leakage. The model copies phrases from examples into outputs (especially in generation tasks). Use varied examples and say "these illustrate style; do not reuse their content".
  • Too many examples. Beyond a point, returns diminish and cost rises. Test whether example 6 actually improves results.

Zero-shot first, then add examples

A good workflow: start zero-shot with crisp instructions, measure, then add examples targeted at the observed errors. This tells you what each example is buying you. Many teams add ten examples on day one and never learn that two of them were doing all the work.

Hands-on: dynamic few-shot selection

For large labelled libraries, retrieve the most similar examples for each input and always include a fixed set of boundary cases. This version uses a simple, dependency-free similarity for clarity; in production, use an embedding model.

import re
from collections import Counter
from math import sqrt

LIBRARY = [  # (input, label) pairs reviewed by a domain expert
    ("We're launching in Riyadh in March and need a paid social partner. Budget approved.", "HOT"),
    ("hi, do u do tiktok ads? thinking about it for next yr", "WARM"),
    ("I'm writing my dissertation on influencer marketing, could I interview you?", "COLD"),
    ("Congratulations! Claim your free SEO audit now", "SPAM"),
    ("Our Dubai store needs this live before Eid, can you quote?", "HOT"),
    # ... hundreds more
]
BOUNDARY = [LIBRARY[1], LIBRARY[3]]  # always included

def vec(text):
    return Counter(re.findall(r"[a-z0-9]+", text.lower()))

def cosine(a, b):
    common = set(a) & set(b)
    num = sum(a[t] * b[t] for t in common)
    den = sqrt(sum(v * v for v in a.values())) * sqrt(sum(v * v for v in b.values()))
    return num / den if den else 0.0

def select_examples(message, k=4):
    q = vec(message)
    ranked = sorted(LIBRARY, key=lambda ex: cosine(q, vec(ex[0])), reverse=True)
    chosen = [ex for ex in ranked if ex not in BOUNDARY][:k]
    return chosen + BOUNDARY

def build_prompt(message):
    shots = "\n".join(
        f"<example>\n<input>{i}</input>\n<output>{o}</output>\n</example>"
        for i, o in select_examples(message)
    )
    return f"<examples>\n{shots}\n</examples>\n\n<input>{message}</input>"

Check the label mix of retrieved examples: if all four share one label, the model may be biased toward it. The fixed boundary set counters that.

Examples, structured outputs and reasoning models

Two current patterns work well together:

  • Examples for judgement, schemas for shape. When you use structured outputs (next lesson), the schema guarantees the format, so your examples can focus on hard decisions rather than on showing JSON layout.
  • Examples with reasoning models. Reasoning models still benefit from examples, especially for domain conventions and label boundaries. Keep them diverse; a single worked path can narrow the model's own exploration. If you include example rationales, keep them short.

How to measure an example's value

Run an ablation: evaluate the prompt with all examples, then remove one example at a time and re-run. Examples whose removal does not change accuracy are costing tokens for nothing. Examples whose removal hurts a specific tag (for instance "WARM vs COLD") are doing real work; keep them and note why in your changelog.

Going further

For large example libraries, select examples dynamically: embed your library, and for each new input retrieve the most similar labelled examples to include. This keeps prompts short while making examples maximally relevant. Watch for the risk that retrieved examples all share the same label, which can bias the model; mixing in a fixed set of boundary examples helps.

Key takeaways

  • Models generalise from everything your examples share, including unintended patterns like length or topic.
  • Good examples are relevant, diverse, correct and clearly delimited from the real input.
  • The most valuable examples sit on decision boundaries; mine them from real disagreements.
  • Start zero-shot, measure, then add examples aimed at specific errors, and never reuse test cases as examples.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. All five examples in a summarisation prompt are exactly two sentences long. What will most likely happen?
  2. Where do the most instructive new examples usually come from?
  3. Why include a short reasoning field in examples for judgement tasks?

Put it into practice

Take a classification task you run. Collect 30 real inputs, label them, find the 3 hardest disagreements, and turn them into boundary examples. Measure accuracy before and after on a separate set.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.