Advanced Prompt EngineeringFew-shot examples and structured outputs · Lesson 3 of 17
Few-shot and example design
Video lecture
Few-shot and example design
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Few-shot and example design
A copywriting prompt had four beautiful examples, all about fintech apps. Then a fashion client arrived, and every caption came out sounding like a banking app. The examples worked, just not the way anyone intended. In this lecture you will learn why examples are so powerful, the four properties of good examples, how to find boundary cases, how to select examples dynamically, and how to measure what each example is really buying you.
0:32 The key mental model
Instructions describe what you want; examples show it. For tone, format, edge cases and classification boundaries, a few well-chosen examples often do more than a page of rules. Provider guidance commonly suggests roughly three to five diverse, relevant examples as a starting point. But here is the mental model that matters: the model generalises from every pattern your examples share, including patterns you did not intend. If every example is three sentences, you have set a length rule. If every example is about software, the outputs drift toward software framing.
1:11 Four properties
Good examples have four properties. Relevant: they look like your real inputs, including typos, mixed languages and rambling messages. Diverse: they cover the main categories and at least one tricky boundary, with varied length and phrasing. Correct: one wrong label teaches the wrong rule, so have a domain expert check them. And clearly separated: wrap them in example tags so the model never confuses an example with the real input.
1:41 Worked example: lead intent
Take a lead classifier with four labels: hot means a named budget or timeline within ninety days, warm means a clear need but no budget or timeline, cold means general curiosity, students or job seekers, and spam. One example says: we are launching in Riyadh in March and need a paid social partner, budget approved. That is hot. Another says, in lowercase with no punctuation: hi, do you do TikTok ads, thinking about it for next year. That is warm. That informal second example teaches the warm boundary more effectively than another sentence in the definition ever could.
2:24 Mining boundary examples
The most valuable examples sit on decision boundaries: the message that looks hot but has no timeline, or the complaint that looks like spam but is a real customer. To find them, run your prompt on fifty to one hundred real inputs, collect the cases where the model and a human disagree, and promote the most instructive disagreements into examples. This is a loop, and your examples should evolve with your evaluation set. But never copy test cases into the prompt, or your accuracy numbers will measure memorisation.
3:02 Example 1: product URL slugs
Here is a simple worked example. You want an assistant to turn product names into friendly URL slugs. The instruction says lowercase with hyphens. But what about ampersands, accents or numbers? Two examples settle it instantly. Input: Salt and Pepper Grinder, two pieces. Output: salt-pepper-grinder-2pc. Input: Crème Brûlée Set. Output: creme-brulee-set. From those two, the model learns to drop words like and, convert numbers, abbreviate pieces, and strip accents, rules you might have taken a paragraph to write. That is the power of examples: they carry decisions you did not have to spell out.
3:43 Example 2: lead classifier (illustrative)
Now a business scenario, with illustrative numbers. A lead-routing classifier for a Karachi agency sorts about two thousand inbound messages a month into hot, warm, cold and spam. Zero-shot accuracy on a labelled set of two hundred messages is around seventy-eight percent. The team looks at the disagreements. Most errors are warm messages mislabelled as hot, because they mention a budget without a timeline, and Urdu-English messages mislabelled as spam. They add four boundary examples: two budget-no-timeline warm leads, and two genuine Roman Urdu enquiries. Accuracy on a separate held-out set rises to around eighty-nine percent. Then they run an ablation, removing one example at a time, and discover a fifth example they had added earlier was doing nothing, so they delete it. Numbers illustrative; the method is the lesson.
4:40 Failure modes
Watch for four failure modes. Overfitting to surface features: all hot examples mention a currency, so any currency looks hot. Fix it with contrasting examples. Label imbalance: three spam examples and one hot can bias toward spam. Example leakage, where phrases from examples get copied into outputs; say these illustrate style, do not reuse their content. And too many examples: beyond a point, returns diminish and cost rises. A good workflow is zero-shot first, measure, then add examples targeted at the errors you actually observed.
5:17 Dynamic example selection
For large libraries, select examples dynamically. For each new input, retrieve the most similar labelled examples, and always add a fixed set of boundary cases. The lesson's hands-on code shows this with a simple similarity function you can swap for an embedding model. One caution: if all the retrieved examples share one label, the model may be biased toward it. The fixed boundary set counters that. And when you use structured outputs, the schema guarantees the format, so examples can focus on hard judgements rather than showing JSON layout.
5:56 Examples with reasoning models
Reasoning models still benefit from examples, especially for domain conventions and label boundaries. Keep them diverse, because a single worked path can narrow the model's own exploration. If you include short rationales, like implied timeline within ninety days and an explicit request for a quote, keep them brief. They teach the way of deciding, not just the answer. If your pipeline needs only the label, ask for reasoning in a separate field or tag and parse out the output.
6:30 Measure each example
How do you know an example earns its place? Run an ablation. Evaluate with all examples, then remove one at a time and re-run. If removing an example changes nothing, it is costing tokens for nothing. If removing it hurts a specific tag, like warm versus cold, it is doing real work. Keep it, and note why in your changelog. Many teams add ten examples on day one and never learn that two of them were doing all the work.
7:05 Recap
To recap. Models copy every pattern your examples share. Make examples relevant, diverse, correct and clearly separated. Mine boundary examples from real disagreements, select dynamically for large libraries, and measure each example's value with ablations. Never put test cases in the prompt. Try this now: take a classification task, collect thirty real inputs, find the three hardest disagreements and turn them into boundary examples, then measure accuracy before and after on a separate set. Next: structured outputs and JSON schemas.
Why examples are so powerful
Instructions describe what you want; examples show it. For tone, format, edge-case handling and classification boundaries, a few well-chosen examples often do more than a page of rules. Provider guidance commonly suggests a small handful (roughly three to five) of diverse, relevant examples as a strong starting point.
The mental model: the model generalises from the pattern your examples share, including patterns you did not intend. If all your examples are about software companies, outputs will drift toward software framing. If all examples are three sentences long, you have just set a length rule.
The four properties of good examples
- Relevant. They look like your real inputs, including messiness: typos, mixed languages, long rambling messages.
- Diverse. They cover the main categories and at least one tricky boundary case. Vary length, topic and phrasing so the model does not latch onto surface features.
- Correct. A single wrong label in your examples teaches the wrong rule. Have a domain expert check them.
- Clearly separated. Wrap them in tags such as
<example>so the model does not confuse examples with the actual input.
A worked example: classifying lead intent
<instructions>
Classify each inbound message as HOT, WARM, COLD or SPAM.
HOT = named budget or timeline within 90 days.
WARM = clear need, no budget or timeline.
COLD = general curiosity, students, job seekers.
Return only the label.
</instructions>
<examples>
<example>
<input>We're launching in Riyadh in March and need a paid social partner. Budget approved.</input>
<output>HOT</output>
</example>
<example>
<input>hi, do u do tiktok ads? thinking about it for next yr</input>
<output>WARM</output>
</example>
<example>
<input>I'm writing my dissertation on influencer marketing, could I interview you?</input>
<output>COLD</output>
</example>
<example>
<input>Congratulations! Claim your free SEO audit now at this link</input>
<output>SPAM</output>
</example>
</examples>
<input>{{message}}</input>Notice the second example: informal, lowercase, vague timeline. It teaches the WARM boundary more effectively than another sentence in the definition.
Designing boundary examples
The most valuable examples sit on decision boundaries: the message that looks HOT but lacks a timeline, or the complaint that looks like spam but is a real customer. To find them:
- Run your prompt on 50 to 100 real inputs.
- Collect the cases where the model and a human disagree.
- Promote the most instructive disagreements into examples, with the correct label.
This is a loop, not a one-time task. Your example set should evolve with your evaluation set (covered in the evaluation module), but never copy your test cases into the prompt, or you will fool yourself about accuracy.
Examples with reasoning
For judgement-heavy tasks, include a short rationale in each example:
<example>
<input>Our Dubai store needs this by Eid, can you quote?</input>
<reasoning>Implied timeline within 90 days and explicit request for a quote.</reasoning>
<output>HOT</output>
</example>This teaches the way of deciding, not only the answer. With reasoning-capable models, showing the reasoning style in examples tends to carry over to how the model thinks about new cases. If your pipeline needs just the label, ask for the reasoning in a separate tag and parse out the output.
Failure modes
- Overfitting to surface features. All positive examples mention a currency, so the model treats any currency mention as HOT. Fix with contrasting examples.
- Label imbalance. Three HOT examples and one COLD can bias toward HOT. Roughly balance, or match your real distribution deliberately.
- Example leakage. The model copies phrases from examples into outputs (especially in generation tasks). Use varied examples and say "these illustrate style; do not reuse their content".
- Too many examples. Beyond a point, returns diminish and cost rises. Test whether example 6 actually improves results.
Zero-shot first, then add examples
A good workflow: start zero-shot with crisp instructions, measure, then add examples targeted at the observed errors. This tells you what each example is buying you. Many teams add ten examples on day one and never learn that two of them were doing all the work.
Hands-on: dynamic few-shot selection
For large labelled libraries, retrieve the most similar examples for each input and always include a fixed set of boundary cases. This version uses a simple, dependency-free similarity for clarity; in production, use an embedding model.
import re
from collections import Counter
from math import sqrt
LIBRARY = [ # (input, label) pairs reviewed by a domain expert
("We're launching in Riyadh in March and need a paid social partner. Budget approved.", "HOT"),
("hi, do u do tiktok ads? thinking about it for next yr", "WARM"),
("I'm writing my dissertation on influencer marketing, could I interview you?", "COLD"),
("Congratulations! Claim your free SEO audit now", "SPAM"),
("Our Dubai store needs this live before Eid, can you quote?", "HOT"),
# ... hundreds more
]
BOUNDARY = [LIBRARY[1], LIBRARY[3]] # always included
def vec(text):
return Counter(re.findall(r"[a-z0-9]+", text.lower()))
def cosine(a, b):
common = set(a) & set(b)
num = sum(a[t] * b[t] for t in common)
den = sqrt(sum(v * v for v in a.values())) * sqrt(sum(v * v for v in b.values()))
return num / den if den else 0.0
def select_examples(message, k=4):
q = vec(message)
ranked = sorted(LIBRARY, key=lambda ex: cosine(q, vec(ex[0])), reverse=True)
chosen = [ex for ex in ranked if ex not in BOUNDARY][:k]
return chosen + BOUNDARY
def build_prompt(message):
shots = "\n".join(
f"<example>\n<input>{i}</input>\n<output>{o}</output>\n</example>"
for i, o in select_examples(message)
)
return f"<examples>\n{shots}\n</examples>\n\n<input>{message}</input>"Check the label mix of retrieved examples: if all four share one label, the model may be biased toward it. The fixed boundary set counters that.
Examples, structured outputs and reasoning models
Two current patterns work well together:
- Examples for judgement, schemas for shape. When you use structured outputs (next lesson), the schema guarantees the format, so your examples can focus on hard decisions rather than on showing JSON layout.
- Examples with reasoning models. Reasoning models still benefit from examples, especially for domain conventions and label boundaries. Keep them diverse; a single worked path can narrow the model's own exploration. If you include example rationales, keep them short.
How to measure an example's value
Run an ablation: evaluate the prompt with all examples, then remove one example at a time and re-run. Examples whose removal does not change accuracy are costing tokens for nothing. Examples whose removal hurts a specific tag (for instance "WARM vs COLD") are doing real work; keep them and note why in your changelog.
Going further
For large example libraries, select examples dynamically: embed your library, and for each new input retrieve the most similar labelled examples to include. This keeps prompts short while making examples maximally relevant. Watch for the risk that retrieved examples all share the same label, which can bias the model; mixing in a fixed set of boundary examples helps.
Key takeaways
- Models generalise from everything your examples share, including unintended patterns like length or topic.
- Good examples are relevant, diverse, correct and clearly delimited from the real input.
- The most valuable examples sit on decision boundaries; mine them from real disagreements.
- Start zero-shot, measure, then add examples aimed at specific errors, and never reuse test cases as examples.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Take a classification task you run. Collect 30 real inputs, label them, find the 3 hardest disagreements, and turn them into boundary examples. Measure accuracy before and after on a separate set.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.