Advanced Prompt EngineeringSystem prompts and context engineering · Lesson 1 of 17
Designing system prompts and roles
Video lecture
Designing system prompts and roles
Chapters
Transcript of the narration, chapter by chapter — the chapter playing now is highlighted.
0:00 Designing system prompts and roles
A support assistant for a small agency quoted a client a price that did not exist. Nobody had told it not to. The system prompt said: you are a helpful assistant, be professional, do not make things up. In this lecture you will learn to write system prompts that actually prevent problems like that: the five layers of a strong brief, when roles help and when they hurt, why calm reasons beat capital letters, and how to treat a system prompt like versioned code. By the end, you will be able to rewrite any weak system prompt into one you can test and defend.
0:45 The mental model
Start with the right mental model. A system prompt is not a spell. It is the standing brief you would give a highly capable colleague on their first day: who they serve, what good work looks like, what they must never do, and what to do when unsure. Models are strong generalists with zero knowledge of your business. Every gap you leave, they fill with a plausible average. So here is a test you can run on any system prompt: if you handed it to a smart contractor with no other context, could they do the job? If they would need to ask five questions first, the model needs those answers too.
1:34 Five layers
Strong system prompts have five layers. One, role and audience, tied to a real situation: the support assistant for a B2B invoicing product, whose readers are busy, often frustrated finance managers. Two, the objective and success criteria: resolve the question in one reply where possible, otherwise collect the three details a human agent needs. Three, behavioural rules with reasons. Four, knowledge and tools: what reference material and tools exist, and when to use them. Five, the output contract: format, length, tone, language, and what to do when information is missing.
2:13 Rules that generalise
Layer three deserves extra attention: rules with reasons. Compare never quote prices, with never quote prices, because pricing varies by contract and wrong quotes create legal disputes. The reason lets the model generalise to cases you did not list, like a customer asking what a typical project costs. Also, say what to do, not only what to avoid. Write in flowing paragraphs beats do not use bullet points. And quantify vague words. Short means different things to different people; under one hundred and twenty words does not.
2:51 Example 1: a homework tutor
Let us start with a simple worked example. A teacher wants an assistant that helps eleven-year-olds with maths homework. The weak system prompt: you are a helpful maths tutor. The strong one, built from the five layers. Role and audience: you help eleven-year-old students with homework; they get discouraged easily. Objective: the student should reach the answer themselves. Rules with reasons: never give the final answer directly, because the goal is learning; if the student is stuck after two hints, show a similar worked example instead. Knowledge: the curriculum topics for this term. Output: short sentences, one question at a time, friendly tone. Now run the same student question through both. The first gives the answer. The second asks: what do you think we should do with the brackets first? That single difference is the whole point of the lesson.
3:52 Example 2: a support assistant (illustrative)
Now a realistic business scenario. The numbers are illustrative. A software company in Dubai runs a support assistant handling about three thousand conversations a month. Too many were escalated to humans, around forty percent, because the assistant kept saying please contact support for anything slightly unusual. The team rewrote the system prompt. They added an explicit objective: resolve in one reply where possible; otherwise collect account email, plan and error message before escalating. They added reasons to rules, for example: do not discuss refunds, because refund decisions need account history the assistant cannot see. And they added an escape hatch with a structured escalation note. After testing on two hundred past conversations, escalations in the test set dropped to around twenty-five percent, and the escalations that remained arrived with the details agents needed. Same model. Better brief.
4:52 Roles: benefits and failure modes
Now roles. Role prompting works because it sets defaults for vocabulary, depth and priorities. A senior data engineer reviewing a pipeline for reliability gives a different review than a helpful assistant. But roles fail in three ways. Persona theatre, like an elaborate backstory about Max who loves hiking, burns tokens and rarely improves quality. Overclaiming authority, like you are a licensed doctor, can push the model to sound more certain than it should. And conflicting roles, a ruthless critic who is also warm and encouraging, produce mush. Pick one primary stance and say how to handle the tension.
5:35 Newest models: more literal
Here is something that changed with the newest models. They follow instructions more literally and more completely than their predecessors. That has three consequences. Say everything you want, because they are less likely to guess you wanted extra effort. Remove legacy scaffolding, such as always double-check or you must use the search tool, which was written for older models and now causes over-triggering: excessive tool calls, hedging, even refusals. And drop old tricks that no longer apply. For example, several current APIs, including recent Claude models, no longer accept pre-filling the start of the assistant's reply. Use structured outputs or clear format instructions instead.
6:20 Worked example: lead replies
Let us fix the agency example. The weak version: you are a helpful assistant for our agency, be professional, do not make things up. The strong version: you draft first-pass replies to inbound leads for a small digital agency serving e-commerce brands in the UK and the Gulf. The goal is a reply the account manager can send after light edits, which acknowledges the lead's situation, asks at most two qualifying questions and proposes a twenty-minute call. Rules: mention only services in the services list; do not quote prices or timelines, because they depend on scoping; match the lead's language, English or Arabic. If the message is spam, reply with the single word SKIP. Output: plain text, eighty to one hundred and fifty words.
7:14 System vs user turn
Where does each piece of content belong? Put stable, reusable instructions, the role, rules and format, in the system prompt. Put task-specific material, this lead's email or today's data, in the user turn, wrapped in clear tags. This separation makes prompts easier to version, cache and test, and it keeps untrusted content out of the highest-authority position. In the lesson's hands-on code, the system prompt is a template file with named variables, rendered once and kept byte-identical between requests, and the lead's message arrives inside a lead message tag.
7:53 System prompts as code
Finally, treat the system prompt as code. Store it in version control, template variables explicitly, and keep a changelog of why each rule exists. A rule with no recorded reason is a rule nobody will dare delete. And watch three failure modes: instruction overload, where forty rules make the model worse at all of them; hidden contradictions, like always be concise in one section and explain thoroughly in another; and stale rules that fixed an old model's quirk but now over-correct. Before shipping, run the prompt on twenty to fifty realistic messages, including out-of-scope, rude, Arabic and spam examples.
8:36 Recap
Let us recap. A system prompt is a day-one brief with five layers: role and audience, objective, rules with reasons, knowledge and tools, and an output contract. Roles help when tied to a real situation. Calm, specific, reasoned instructions beat capital letters, especially on the newest models. Keep stable instructions in the system prompt and task material in the user turn. Try this now: take a real system prompt, rewrite it with the five layers, add a reason to every rule, and compare ten outputs before and after. Next, we go beyond the prompt, to everything the model sees: context engineering.
The mental model: a brief for a brilliant new colleague
A system prompt is not a magic spell. It is the standing brief you would give a highly capable colleague on their first day: who they are serving, what good work looks like, what they must never do, and what to do when they are unsure. Models are strong generalists with no knowledge of your business, your customers or your taste. Every gap you leave, the model fills with a plausible average. Advanced prompt engineering is mostly the discipline of removing those gaps deliberately.
A useful test: if you handed your system prompt to a smart contractor with no other context, could they do the job well? If they would need to ask five questions first, the model needs those answers too.
The five layers of a strong system prompt
- Role and audience. Not "You are an expert marketer" (which adds little on its own) but a role tied to a situation: "You are the support assistant for a B2B invoicing product. Your readers are finance managers at small companies who are busy and often frustrated."
- Objective and success criteria. What outcome matters? "Resolve the question in one reply where possible; if not, collect the three details a human agent needs."
- Behavioural rules with reasons. "Never quote prices, because pricing varies by contract and wrong quotes create legal disputes." Explaining why lets the model generalise to cases you did not list.
- Knowledge and tools. What reference material is available, which tools exist and when to use them.
- Output contract. Format, length, tone, language and what to do when information is missing.
Role design: when personas help and when they hurt
Role prompting works because it sets defaults for vocabulary, depth and priorities. "You are a senior data engineer reviewing a pipeline for reliability" produces a different review than "You are a helpful assistant." But roles have failure modes:
- Persona theatre. Elaborate backstories ("You are Max, 47, who loves hiking") consume tokens and rarely improve task quality.
- Overclaiming authority. "You are a licensed doctor" can push the model to sound more certain than it should. Prefer "You help clinicians by summarising literature; you do not diagnose."
- Conflicting roles. "Be a ruthless critic" plus "be warm and encouraging" produces mush. Pick one primary stance and specify how to handle the tension: "Be candid about problems, framed constructively."
Instructions: specific, positive, reasoned
Models follow instructions more literally than many people expect, especially newer ones. Three rules of thumb:
- Say what to do, not only what to avoid. "Write in flowing paragraphs" beats "Don't use bullet points."
- Quantify vague words. "Short" means different things to different people; "under 120 words" does not.
- Give the reason. "Avoid jargon because readers are first-time founders" lets the model judge edge cases.
Avoid SHOUTING and stacked "CRITICAL!!!" markers. With capable modern models, aggressive emphasis tends to cause overreaction, such as refusing reasonable requests near a forbidden topic. Calm, precise language with reasons works better.
A worked example
Weak:
You are a helpful assistant for our agency. Be professional. Don't make things up.Strong:
You draft first-pass replies to inbound leads for a small digital agency
that serves e-commerce brands in the UK and the Gulf.
Goal: a reply the account manager can send after light edits, which
(1) acknowledges the lead's specific situation, (2) asks at most two
qualifying questions, and (3) proposes a 20-minute call.
Rules:
- Only mention services listed in the services section; if the lead asks
for something else, say we will confirm whether it is a fit.
- Do not quote prices or timelines; these depend on scoping.
- Match the lead's language (English or Arabic) and formality.
- If the message is spam or unrelated, reply with the single word SKIP.
Output: plain text email body, 80-150 words, no subject line.Notice what changed: a real audience, a measurable goal, rules with reasons, an explicit escape hatch (SKIP), and an output contract.
System prompt vs user turn
Put stable, reusable instructions (role, rules, format) in the system prompt. Put the task-specific material (this lead's email, today's data) in the user turn. This separation makes prompts easier to version, cache and test, and it keeps untrusted content out of the highest-authority position.
Failure modes to watch
- Instruction overload. Past a certain point, adding rules makes the model worse at all of them. If you have forty rules, group them, delete the ones the model already follows, and test.
- Hidden contradictions. "Always be concise" in one section and "explain thoroughly" in another. Read your prompt as a hostile reviewer.
- Stale rules. Rules written to fix an old model's quirk may now cause over-correction. Re-test rules when you change models.
What changed with the newest models
Frontier models released through 2025 and 2026 follow instructions more literally and more completely than their predecessors. Three practical consequences:
- Say everything you want. If you want the assistant to go beyond the obvious (suggest alternatives, add edge cases), ask for it. Newer models tend not to guess that you wanted more.
- Remove legacy scaffolding. Rules such as "ALWAYS double-check", "think step by step before every answer" or "you MUST use the search tool" were written for older models. On current models they often cause over-triggering: excessive tool calls, verbose hedging or refusals. Rewrite them calmly with the reason, or delete them and re-test.
- Some old tricks no longer apply. For example, several current APIs, including recent Claude models, no longer accept a pre-filled start to the assistant's reply ("prefill"). Use structured outputs or clear format instructions instead.
Provider migration guides are the best source for these shifts; read them whenever you change model generation.
Hands-on: a system prompt as a versioned template
Store the system prompt as a file with explicit variables, render it in code, and keep the stable part byte-identical between requests so it can be cached (see the prompt caching lesson).
# pip install anthropic
import os
from string import Template
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5") # check the current models page
SYSTEM = Template("""You draft first-pass replies to inbound leads for $agency, a small
digital agency serving e-commerce brands in the UK and the Gulf.
<goal>
A reply the account manager can send after light edits: acknowledge the lead's
situation, ask at most two qualifying questions, propose a 20-minute call.
</goal>
<rules>
- Mention only services in <services>; if asked for anything else, say we will
confirm fit, because we lose trust when we over-promise.
- Do not quote prices or timelines; they depend on scoping.
- Match the lead's language (English or Arabic) and formality.
- If the message is spam or unrelated, reply with exactly: SKIP
</rules>
<services>$services</services>
<output>Plain-text email body, 80-150 words, no subject line.</output>""")
system_prompt = SYSTEM.substitute(agency="Northwind Digital",
services="Paid social; Shopify CRO; email automation")
def draft_reply(lead_message: str) -> str:
try:
resp = client.messages.create(
model=MODEL,
max_tokens=1024,
system=system_prompt,
messages=[{"role": "user", "content": f"<lead_message>\n{lead_message}\n</lead_message>"}],
)
except anthropic.APIStatusError as err:
raise RuntimeError(f"API error {err.status_code}") from err
return "".join(b.text for b in resp.content if b.type == "text").strip()
print(draft_reply("Salam, we launch our abaya store in Riyadh in March. Can you run TikTok ads?"))The same structure ports to other providers: in the OpenAI Responses API the system-level brief goes in instructions; in the Gemini API it goes in the system instruction field. Keep the brief provider-neutral where you can and adapt only the call.
How to test a system prompt
Before shipping, run it against 20 to 50 realistic messages, including out-of-scope, Arabic, rude and spam examples, and check each rule is followed. When you edit a rule, re-run the whole set: a fix for one behaviour frequently breaks another.
Going further
If you work through an API, treat the system prompt as code: store it in version control, template variables explicitly (for example {{customer_tier}}), and keep a changelog of why each rule exists. A rule with no recorded reason is a rule nobody will dare to delete.
Key takeaways
- A system prompt is a job brief: role, objective, rules with reasons, knowledge/tools and an output contract.
- Roles help when tied to a real situation and audience; elaborate personas rarely add quality.
- Positive, quantified, reasoned instructions beat vague prohibitions and ALL-CAPS emphasis.
- Keep stable instructions in the system prompt and task-specific or untrusted content in the user turn.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Take a real system prompt you use (or write one for a support assistant). Rewrite it using the five layers, add a reason to every rule, and compare 10 outputs before and after.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.