Latest AI Techniques: RAG, Tool Use, Agents & MCPAgent loops, multi-agent patterns and memory · Lesson 12 of 20

Context engineering: curating what the model sees

Article · 14 min · 9 min lecture

Video lecture

Context engineering: curating what the model sees

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Context engineering

  • What the model sees decides what it does
  • Curate, don't dump
  • Quality, cost and speed

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From prompt engineering to context engineering

Prompt engineering asks: "What words should I write?" Context engineering asks a bigger question: "Of everything this model could see right now (instructions, conversation, retrieved documents, tool definitions, tool results, memories, examples), what is the smallest set of high-signal information that makes the right next step likely?" As systems moved from single prompts to agents that run for dozens of steps, that second question became the main determinant of quality, cost and latency.

Context windows are now very large (frontier models from several vendors accept around a million tokens at the time of writing; check current limits). Large windows do not make the question go away. More context costs more, slows responses and, past a point, degrades attention: relevant details get lost among irrelevant ones, a problem practitioners call context rot. Treat context as a scarce budget, not a bin.

The anatomy of a context window

ComponentTypical contentEngineering questions
System instructionsRole, goals, rules, output formatRight altitude? Stable enough to cache?
Tool definitionsNames, descriptions, schemasHow many? Loaded all at once or on demand?
ExamplesA few diverse, canonical casesDo they cover edge cases without bloating?
Retrieved knowledgePassages, recordsHow many, how ranked, how fresh?
Conversation / trajectoryMessages, tool calls, tool resultsWhat can be cleared, summarised or moved out?
MemoryUser and task factsWhat is relevant now?

Principles that hold up

1. Right altitude for instructions. Avoid both extremes: brittle if-this-then-that rules for every case, and vague aspirations ("be helpful"). Give clear goals, priorities, constraints and a few heuristics. Organise with headings or XML-style tags so the model can find things.

2. Just-in-time context. Rather than loading everything up front, give the agent lightweight references (file paths, record IDs, search tools) and let it fetch details when needed. This mirrors how people work: we do not memorise the whole drive; we know where to look.

3. Tool results are context too. A single verbose API response can dominate the window. Return concise, structured results with a detail_level option; paginate; summarise large payloads before they enter the transcript.

4. Clear, compact, or offload as tasks grow.

  • Clearing: remove old tool results that have already been used (some APIs now offer automatic context editing that clears stale tool results).
  • Compaction: replace older history with a structured summary of decisions, facts, open issues and next steps (some APIs offer server-side compaction for long conversations).
  • Structured note-taking: the agent writes progress notes or a to-do file outside the window and reads them back, which survives even a full context reset.
  • Sub-agents: delegate reading-heavy sub-tasks to agents with fresh context that return only condensed findings.

5. Keep the stable prefix stable. Prompt caching works on exact prefixes: tools, then system prompt, then messages. Put fixed instructions and tool definitions first and never inject timestamps or per-request IDs into them; put volatile content last. A cache-friendly layout can cut cost and latency dramatically for repeated calls.

6. Separate trusted from untrusted. Label retrieved documents, emails and web pages as data (for example inside <document> tags) and state in the instructions that content inside them is never a command.

Worked example: an e-commerce operations agent

A Lahore-based online retailer runs an agent that investigates delayed orders each morning: it reads a report of late orders, checks courier tracking, reviews support tickets and drafts customer updates for approval.

Version 1 loaded 40 tools, the full courier API responses and every ticket's full history. Runs were slow, costly, and the agent frequently lost track of which orders it had handled.

Version 2 applied context engineering:

  • Tools cut to 7 task-level tools; courier lookups return five fields instead of the raw payload.
  • A progress.md note lists orders handled, findings and remaining work; the agent updates it every five orders.
  • Old tool results are cleared once summarised into the note.
  • Ticket histories are fetched only for orders flagged as "customer contacted us".
  • The stable system prompt and tool list come first, so every run reuses the cache.

Median run cost and time fell substantially, and the "lost track" failure disappeared from the evaluation suite.

Hands-on: a context budget report

Before optimising, measure. Anthropic's API offers a token-counting endpoint; most providers have an equivalent. This script shows where your tokens go.

import json, os
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")

def count(system="", tools=None, messages=None):
    msgs = messages or [{"role": "user", "content": "."}]
    kwargs = {"model": MODEL, "system": system, "messages": msgs}
    if tools:
        kwargs["tools"] = tools
    return client.messages.count_tokens(**kwargs).input_tokens

def budget_report(system, tools, examples, retrieved, history):
    base = count()
    parts = {
        "system": count(system=system) - base,
        "tools": count(tools=tools) - base,
        "examples": count(messages=[{"role": "user", "content": examples}]) - base,
        "retrieved": count(messages=[{"role": "user", "content": retrieved}]) - base,
        "history": count(messages=history) - base if history else 0,
    }
    total = sum(parts.values())
    for name, n in sorted(parts.items(), key=lambda x: -x[1]):
        print(f"{name:10} {n:8,} tokens  {n / max(total, 1):6.1%}")
    return parts

Run it on a real request from your system. The biggest line is where to start: usually tool results in history, retrieved passages or tool definitions. Then change one thing, re-run your evaluation suite, and confirm quality held while tokens fell.

Pitfalls

  • Assuming a bigger window means you can stop curating.
  • Timestamps or random IDs at the top of the system prompt, which silently break caching.
  • Letting raw tool payloads accumulate for the whole run.
  • Compaction summaries that drop decisions or constraints; test them like any other component.

Key takeaways

  • Context engineering chooses the smallest high-signal set of information for each model call.
  • Large windows do not remove the need to curate: cost, latency and context rot all grow with clutter.
  • Use right-altitude instructions, just-in-time retrieval, concise tool results, clearing, compaction, notes and sub-agents.
  • Keep stable content first for prompt caching, label untrusted content as data, and measure token budgets before optimising.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent’s runs get slower and it forgets which orders it already handled. Which change best addresses both?
  2. Why can a timestamp at the start of the system prompt increase costs?
  3. What does "just-in-time context" mean?

Put it into practice

Run the budget report (or estimate by hand) on one real request in your system. Identify the largest component and propose one change to shrink it without losing quality.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.