Advanced Prompt EngineeringSystem prompts and context engineering · Lesson 2 of 17

Context engineering: what the model sees, and in what order

Article · 13 min · 8 min lecture

Video lecture

Context engineering: what the model sees, and in what order

12 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 12

Context engineering

  • Curate everything the model sees
  • Include, exclude, compress
  • Order and label
  • Long-context and agent strategies

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From prompt engineering to context engineering

As applications grew from single prompts into assistants and agents, practitioners began talking about context engineering: the discipline of curating everything that lands in the model's context window. That includes the system prompt, retrieved documents, conversation history, tool definitions, tool results, examples and the user's request. The model can only reason about what it sees, and it reasons less reliably when what it sees is noisy, contradictory or bloated.

The core idea: context is a scarce, expensive resource with diminishing returns. Even when a model accepts very long inputs, more tokens mean more cost, more latency and more opportunities for distraction. The goal is the smallest set of high-signal information that makes the correct answer likely.

Include, exclude, compress

For every piece of candidate context, ask three questions:

  1. Does the task need it? A pricing question needs the price list, not the company history.
  2. Is it trustworthy and current? Outdated policies are worse than none, because the model will cite them confidently.
  3. Can it be compressed without losing signal? Replace a 40-message chat history with a structured summary of decisions and open questions.

Common things to exclude: boilerplate (navigation text, legal footers), duplicated documents, irrelevant tool outputs from earlier steps, and "just in case" material. Common things to include that people forget: the audience, the date, the definition of done, and domain terms the model might misread (for example, what "active customer" means in your company).

Ordering and structure

Position matters. Durable guidance from model providers includes:

  • Put long documents first and the question last. For long inputs, placing the material near the top and your instructions and query at the end tends to improve answer quality.
  • Label everything. Wrap each type of content in clear delimiters so the model can tell instructions from data. XML-style tags work well across many models:
<documents>
  <document index="1">
    <source>refund-policy-2026.md</source>
    <document_content>...</document_content>
  </document>
  <document index="2">
    <source>shipping-faq.md</source>
    <document_content>...</document_content>
  </document>
</documents>

<instructions>
Answer the customer's question using only the documents above.
Quote the relevant passage first, then answer.
</instructions>

<question>Can I return a sale item after 20 days?</question>
  • Metadata helps. Source names, dates and authority levels ("official policy" vs "forum post") let the model weigh conflicting information.

Long-context strategies

When material genuinely is large, you have several options, roughly from cheapest to most involved:

  • Quote-then-answer. Ask the model to extract the relevant quotes into a tagged section first, then answer from those quotes. This focuses attention and makes answers checkable.
  • Map-reduce. Split the material, process each chunk with the same prompt (for example "extract every obligation with its clause number"), then combine the partial results in a final call.
  • Retrieval. Instead of sending everything, search for the most relevant passages and send only those. This is the basis of retrieval-augmented generation, covered in depth in our RAG and agents course.
  • Structured memory. For long conversations or agent runs, keep a running notes object (decisions, facts, open items) and pass that rather than the full transcript.

Worked example: a board-pack summariser

A strategy team pastes a 120-page board pack and asks, "What are the risks?" The first attempt returns generic risks. Improved context design:

  1. Remove appendices of raw tables the question does not need.
  2. Add metadata: which section is the CFO report vs the marketing update.
  3. Define "risk" for this board: financial, regulatory, operational, with a materiality threshold.
  4. Ask for quotes with page references before the synthesis.
  5. Put the question and output format at the end.

The result is shorter, specific and auditable, because the model was given a sharper problem, not a bigger one.

Failure modes

  • Context rot. Quality can degrade as context grows, particularly for details buried in the middle of very long inputs. Test with your real lengths rather than assuming.
  • Contradictory sources. If two documents disagree, say which wins ("newer policy overrides older").
  • Instruction leakage from data. Retrieved text can contain imperative sentences ("ignore previous instructions"). Delimiting helps, but see the prompt injection lesson for real defences.

Context engineering for long-running assistants and agents

Context engineering matters most when context accumulates: multi-turn chats and agents that call tools dozens of times. Current platforms offer tools for this, and it helps to know the vocabulary:

  • Compaction: summarising older turns into a compact state so the conversation can continue. Some APIs now offer server-side compaction; you can also do it yourself with a summarisation call.
  • Context editing or clearing: removing stale tool results or old reasoning blocks that no longer matter, while keeping the decisions derived from them.
  • External memory: writing notes, plans and facts to a file or store the agent can read back later, instead of keeping everything in the window.
  • Just-in-time retrieval: giving the model lightweight references (file paths, IDs, search tools) and letting it load details only when needed, rather than pre-loading everything.

The design question is always the same: which tokens earn their place in this particular call?

Hands-on: a context budget report

Measure before you optimise. This script counts tokens per component with the Claude token-counting endpoint, so you can see what dominates.

import os
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")

components = {
    "system": open("prompts/system.md", encoding="utf-8").read(),
    "retrieved_docs": open("context/retrieved.md", encoding="utf-8").read(),
    "history_summary": open("context/history.md", encoding="utf-8").read(),
    "question": "Can I return a sale item after 20 days?",
}

def count(text: str) -> int:
    result = client.messages.count_tokens(
        model=MODEL, messages=[{"role": "user", "content": text}]
    )
    return result.input_tokens

report = {name: count(text) for name, text in components.items()}
total = sum(report.values())
for name, n in sorted(report.items(), key=lambda kv: -kv[1]):
    print(f"{name:16} {n:7} tokens  {100 * n / total:5.1f}%")

Counts include a small per-message overhead, so treat them as close estimates. Token counts differ between model families because tokenisers differ; always count with the model you will use.

Worked example: trimming a support bot

A support team finds that 70% of each request (illustrative figure) is old tool output: full order records returned by a lookup tool on earlier turns. They change the tool to return only the fields the model needs, replace earlier tool results with a one-line summary after each answer, and move stable policy text into a cached prefix. Answers get faster and cheaper, and accuracy on their eval set holds. The lesson: the biggest context wins are usually structural, not wording tweaks.

Checklist before every production prompt

  • Is every component needed for this call?
  • Is anything stale, duplicated or contradictory?
  • Are long documents above the question, labelled with source and date?
  • Are stable parts first, so they can be cached?
  • Do you know the token count of each part?

Going further

Measure context, don't guess. Log token counts per component (system, history, retrieval, tools) for real traffic. You will often find that one component, such as verbose tool results, dominates cost while adding little. Trimming it is frequently the single cheapest quality and latency win available.

Key takeaways

  • Context engineering curates everything in the window: instructions, documents, history, tools and results.
  • Aim for the smallest set of high-signal context; more tokens add cost, latency and distraction.
  • Put long material first, instructions and question last, and label each part with clear delimiters.
  • For large inputs use quote-then-answer, map-reduce, retrieval or structured memory instead of dumping everything.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A model answers questions about a 90-page contract but misses clauses. What is the best first change?
  2. Which item is most worth ADDING to the context of a customer analytics assistant?
  3. Two retrieved policies conflict. What should the prompt do?

Put it into practice

Pick one long-document task you do. Rebuild the prompt with labelled documents at the top, a quote-extraction step, and the question at the end. Compare accuracy on five questions you know the answers to.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.