Latest AI Techniques: RAG, Tool Use, Agents & MCPAgent loops, multi-agent patterns and memory · Lesson 12 of 20
Context engineering: curating what the model sees
Video lecture
Context engineering: curating what the model sees
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Context engineering
Picture a brilliant new colleague on their first morning. You could hand them every document in the company, every email thread and every tool manual, stacked on their desk. Or you could give them a clear brief, the three documents that matter, and tell them where to find the rest. The second colleague does better work, faster. That's context engineering, and for AI agents it has become the single biggest lever on quality, cost and speed.
0:33 Why it matters
Why is this becoming the key skill? Because as agents run longer and use more tools, what goes into each model call is increasingly assembled by software, not typed by a person. Every tool result, every retrieved passage, every memory is a decision. Think of it like packing a small suitcase for a business trip. You could bring your whole wardrobe in a lorry. But you'll move faster, and find things more easily, with exactly what the trip needs.
1:07 From prompts to context
Prompt engineering asks what words to write. Context engineering asks a bigger question: of everything this model could see right now, instructions, conversation, retrieved documents, tool definitions, tool results, memories and examples, what's the smallest set of high-signal information that makes the right next step likely? Context windows are huge now, around a million tokens on some frontier models. But more context costs more, slows responses and, past a point, dilutes attention. Practitioners call that context rot.
1:40 Anatomy of a context window
Let's look inside a context window. System instructions: role, goals, rules and format. Tool definitions. A few examples. Retrieved knowledge. The conversation and the agent's trajectory of tool calls and results. And memory. For each part you ask different questions. Are the instructions at the right altitude, not brittle rules for every case, not vague wishes? Are all forty tools needed on every call? Are those retrieved passages the best five, or just the first twenty?
2:13 Principles
Here are the principles that hold up. Just-in-time context: give the agent references, like file paths, record IDs and search tools, and let it fetch details when needed, the way you know where things live without memorising the whole drive. Remember tool results are context too, so return concise fields, not raw payloads. And as tasks grow, clear, compact or offload. Clear old tool results once used. Compact older history into a summary of decisions, facts and next steps. Have the agent keep notes in a file. Or hand reading-heavy work to sub-agents with fresh context.
2:55 Cache and trust
Two more principles. First, keep the stable prefix stable. Prompt caching matches exact prefixes, tools first, then the system prompt, then messages. Put fixed content first, and never inject timestamps or request IDs into it. Put volatile content last. Second, separate trusted from untrusted. Wrap documents, emails and web pages in clearly labelled tags, and state that anything inside them is data, never a command.
3:23 Simple example: courier lookup
A simple example. A support agent handles a ticket about a late order. Version one pastes the full courier API response, two thousand lines of JSON, into its context. Version two calls the same courier API but returns five fields: status, last scan location, last scan time, estimated delivery and exception code. The answer quality is the same or better, the call is cheaper and faster, and there's far more room left for the rest of the conversation.
3:57 Worked example: delayed-orders agent
Here's a worked example. A Lahore online retailer runs an agent each morning to investigate delayed orders. Version one loaded forty tools, full courier payloads and every ticket's history. It was slow, expensive and kept losing track of which orders it had handled. Version two cut to seven task-level tools, returned five courier fields instead of the raw response, kept a progress note updated every five orders, cleared old results once summarised, fetched ticket histories only when a customer had actually contacted support, and kept the stable prompt first for caching. Cost and time fell sharply, and the losing-track failure vanished from the test suite.
4:43 Business example (illustrative)
Numbers on the Lahore retailer, illustrative. Version one used around a hundred and eighty thousand input tokens per morning run and took about fourteen minutes. Version two used around forty thousand and took about five minutes, with most of the stable prefix served from cache. Their evaluation suite of twenty delayed-order scenarios passed slightly more often in version two, because the agent stopped confusing orders it had already handled.
5:13 Hands-on in the lesson
In the hands-on section you'll build a context budget report. It uses a token-counting endpoint to measure how many tokens your system prompt, tools, examples, retrieved passages and history each consume on a real request, and prints them largest first. Start with the biggest line, change one thing, re-run your evaluation suite, and confirm quality held while tokens fell. Watch out for the classic pitfalls: assuming a big window means no curation, timestamps that break caching, raw payloads piling up, and compaction summaries that quietly drop constraints.
5:51 Common mistakes
Common mistakes. Putting dynamic data, like today's date or a user ID, at the very top of the system prompt, which breaks caching. Loading every tool on every call. Keeping raw tool outputs for the whole run. Writing compaction summaries that lose decisions or constraints. And never measuring token usage by component, so you optimise the wrong thing. Measure, then trim the biggest line.
6:19 How you'll know it's working
How will you know your context engineering worked? Tokens per task fall while your evaluation scores hold or improve. Cache hit rates rise on repeated calls, which you can see in usage data. Latency drops. And long-running agents stop showing symptoms like forgetting earlier decisions or repeating work. If quality drops after a trim, you cut something that mattered, so put it back and trim elsewhere.
6:48 Watch me do it: budget report
Watch me do it. I open the budget report. First, count calls the token counting endpoint with a system prompt, optional tools and messages, and returns input tokens. Next, I measure a baseline with an empty request, so I can subtract it. Then budget report counts each component separately: the system prompt alone, the tools alone, the examples, the retrieved passages and the conversation history. It sorts them largest first and prints each as tokens and a percentage. I run it on a real support request. History is about forty percent, tools twenty-four, retrieved text twenty. So I start with history: I clear tool results older than five turns and replace them with a short progress note. I re-run the report and history drops to about fifteen percent. Then I re-run my evaluation suite to confirm quality held.
7:48 Recap
To recap: context engineering curates the smallest high-signal context for every call. Use right-altitude instructions, just-in-time retrieval and concise tool results. Clear, compact, take notes or delegate as tasks grow. Keep the stable prefix first for caching and label untrusted content as data. Your next step is to measure one real request, find its biggest component, and shrink it without losing quality. Next, we'll compare the frameworks and SDKs you can build agents with.
8:20 Try this now (20 minutes)
Try this now. Take one real request your AI system makes, or one long agent run. Use the budget report from the lesson, or estimate by eye, and write down roughly how much of the context is instructions, tools, examples, retrieved text and history. Circle the biggest. Then write one change that would shrink it, such as fewer tools, trimmed tool results or clearing old history, and test it against your evaluation set.
From prompt engineering to context engineering
Prompt engineering asks: "What words should I write?" Context engineering asks a bigger question: "Of everything this model could see right now (instructions, conversation, retrieved documents, tool definitions, tool results, memories, examples), what is the smallest set of high-signal information that makes the right next step likely?" As systems moved from single prompts to agents that run for dozens of steps, that second question became the main determinant of quality, cost and latency.
Context windows are now very large (frontier models from several vendors accept around a million tokens at the time of writing; check current limits). Large windows do not make the question go away. More context costs more, slows responses and, past a point, degrades attention: relevant details get lost among irrelevant ones, a problem practitioners call context rot. Treat context as a scarce budget, not a bin.
The anatomy of a context window
| Component | Typical content | Engineering questions |
|---|---|---|
| System instructions | Role, goals, rules, output format | Right altitude? Stable enough to cache? |
| Tool definitions | Names, descriptions, schemas | How many? Loaded all at once or on demand? |
| Examples | A few diverse, canonical cases | Do they cover edge cases without bloating? |
| Retrieved knowledge | Passages, records | How many, how ranked, how fresh? |
| Conversation / trajectory | Messages, tool calls, tool results | What can be cleared, summarised or moved out? |
| Memory | User and task facts | What is relevant now? |
Principles that hold up
1. Right altitude for instructions. Avoid both extremes: brittle if-this-then-that rules for every case, and vague aspirations ("be helpful"). Give clear goals, priorities, constraints and a few heuristics. Organise with headings or XML-style tags so the model can find things.
2. Just-in-time context. Rather than loading everything up front, give the agent lightweight references (file paths, record IDs, search tools) and let it fetch details when needed. This mirrors how people work: we do not memorise the whole drive; we know where to look.
3. Tool results are context too. A single verbose API response can dominate the window. Return concise, structured results with a detail_level option; paginate; summarise large payloads before they enter the transcript.
4. Clear, compact, or offload as tasks grow.
- Clearing: remove old tool results that have already been used (some APIs now offer automatic context editing that clears stale tool results).
- Compaction: replace older history with a structured summary of decisions, facts, open issues and next steps (some APIs offer server-side compaction for long conversations).
- Structured note-taking: the agent writes progress notes or a to-do file outside the window and reads them back, which survives even a full context reset.
- Sub-agents: delegate reading-heavy sub-tasks to agents with fresh context that return only condensed findings.
5. Keep the stable prefix stable. Prompt caching works on exact prefixes: tools, then system prompt, then messages. Put fixed instructions and tool definitions first and never inject timestamps or per-request IDs into them; put volatile content last. A cache-friendly layout can cut cost and latency dramatically for repeated calls.
6. Separate trusted from untrusted. Label retrieved documents, emails and web pages as data (for example inside <document> tags) and state in the instructions that content inside them is never a command.
Worked example: an e-commerce operations agent
A Lahore-based online retailer runs an agent that investigates delayed orders each morning: it reads a report of late orders, checks courier tracking, reviews support tickets and drafts customer updates for approval.
Version 1 loaded 40 tools, the full courier API responses and every ticket's full history. Runs were slow, costly, and the agent frequently lost track of which orders it had handled.
Version 2 applied context engineering:
- Tools cut to 7 task-level tools; courier lookups return five fields instead of the raw payload.
- A
progress.mdnote lists orders handled, findings and remaining work; the agent updates it every five orders. - Old tool results are cleared once summarised into the note.
- Ticket histories are fetched only for orders flagged as "customer contacted us".
- The stable system prompt and tool list come first, so every run reuses the cache.
Median run cost and time fell substantially, and the "lost track" failure disappeared from the evaluation suite.
Hands-on: a context budget report
Before optimising, measure. Anthropic's API offers a token-counting endpoint; most providers have an equivalent. This script shows where your tokens go.
import json, os
import anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
def count(system="", tools=None, messages=None):
msgs = messages or [{"role": "user", "content": "."}]
kwargs = {"model": MODEL, "system": system, "messages": msgs}
if tools:
kwargs["tools"] = tools
return client.messages.count_tokens(**kwargs).input_tokens
def budget_report(system, tools, examples, retrieved, history):
base = count()
parts = {
"system": count(system=system) - base,
"tools": count(tools=tools) - base,
"examples": count(messages=[{"role": "user", "content": examples}]) - base,
"retrieved": count(messages=[{"role": "user", "content": retrieved}]) - base,
"history": count(messages=history) - base if history else 0,
}
total = sum(parts.values())
for name, n in sorted(parts.items(), key=lambda x: -x[1]):
print(f"{name:10} {n:8,} tokens {n / max(total, 1):6.1%}")
return partsRun it on a real request from your system. The biggest line is where to start: usually tool results in history, retrieved passages or tool definitions. Then change one thing, re-run your evaluation suite, and confirm quality held while tokens fell.
Pitfalls
- Assuming a bigger window means you can stop curating.
- Timestamps or random IDs at the top of the system prompt, which silently break caching.
- Letting raw tool payloads accumulate for the whole run.
- Compaction summaries that drop decisions or constraints; test them like any other component.
Key takeaways
- Context engineering chooses the smallest high-signal set of information for each model call.
- Large windows do not remove the need to curate: cost, latency and context rot all grow with clutter.
- Use right-altitude instructions, just-in-time retrieval, concise tool results, clearing, compaction, notes and sub-agents.
- Keep stable content first for prompt caching, label untrusted content as data, and measure token budgets before optimising.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run the budget report (or estimate by hand) on one real request in your system. Identify the largest component and propose one change to shrink it without losing quality.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.