---
title: "Context engineering: what the model sees, and in what order"
description: "From prompt engineering to context engineering As applications grew from single prompts into assistants and agents, practitioners began talking about…"
url: https://optimizeall.com/learn/advanced-prompt-engineering/context-engineering
updated: 2026-10-05
---

Advanced Prompt Engineering · System prompts and context engineering · lesson 2 of 17 · 13 min

# Context engineering: what the model sees, and in what order

## From prompt engineering to context engineering

As applications grew from single prompts into assistants and agents, practitioners began talking about *context engineering*: the discipline of curating everything that lands in the model's context window. That includes the system prompt, retrieved documents, conversation history, tool definitions, tool results, examples and the user's request. The model can only reason about what it sees, and it reasons less reliably when what it sees is noisy, contradictory or bloated.

The core idea: **context is a scarce, expensive resource with diminishing returns.** Even when a model accepts very long inputs, more tokens mean more cost, more latency and more opportunities for distraction. The goal is the *smallest set of high-signal information* that makes the correct answer likely.

## Include, exclude, compress

For every piece of candidate context, ask three questions:

1. **Does the task need it?** A pricing question needs the price list, not the company history.
2. **Is it trustworthy and current?** Outdated policies are worse than none, because the model will cite them confidently.
3. **Can it be compressed without losing signal?** Replace a 40-message chat history with a structured summary of decisions and open questions.

Common things to **exclude**: boilerplate (navigation text, legal footers), duplicated documents, irrelevant tool outputs from earlier steps, and "just in case" material. Common things to **include** that people forget: the audience, the date, the definition of done, and domain terms the model might misread (for example, what "active customer" means in your company).

## Ordering and structure

Position matters. Durable guidance from model providers includes:

- **Put long documents first and the question last.** For long inputs, placing the material near the top and your instructions and query at the end tends to improve answer quality.
- **Label everything.** Wrap each type of content in clear delimiters so the model can tell instructions from data. XML-style tags work well across many models:

```xml
<documents>
  <document index="1">
    <source>refund-policy-2026.md</source>
    <document_content>...</document_content>
  </document>
  <document index="2">
    <source>shipping-faq.md</source>
    <document_content>...</document_content>
  </document>
</documents>

<instructions>
Answer the customer's question using only the documents above.
Quote the relevant passage first, then answer.
</instructions>

<question>Can I return a sale item after 20 days?</question>
```

- **Metadata helps.** Source names, dates and authority levels ("official policy" vs "forum post") let the model weigh conflicting information.

## Long-context strategies

When material genuinely is large, you have several options, roughly from cheapest to most involved:

- **Quote-then-answer.** Ask the model to extract the relevant quotes into a tagged section first, then answer from those quotes. This focuses attention and makes answers checkable.
- **Map-reduce.** Split the material, process each chunk with the same prompt (for example "extract every obligation with its clause number"), then combine the partial results in a final call.
- **Retrieval.** Instead of sending everything, search for the most relevant passages and send only those. This is the basis of retrieval-augmented generation, covered in depth in our RAG and agents course.
- **Structured memory.** For long conversations or agent runs, keep a running notes object (decisions, facts, open items) and pass that rather than the full transcript.

## Worked example: a board-pack summariser

A strategy team pastes a 120-page board pack and asks, "What are the risks?" The first attempt returns generic risks. Improved context design:

1. Remove appendices of raw tables the question does not need.
2. Add metadata: which section is the CFO report vs the marketing update.
3. Define "risk" for this board: financial, regulatory, operational, with a materiality threshold.
4. Ask for quotes with page references before the synthesis.
5. Put the question and output format at the end.

The result is shorter, specific and auditable, because the model was given a sharper problem, not a bigger one.

## Failure modes

- **Context rot.** Quality can degrade as context grows, particularly for details buried in the middle of very long inputs. Test with your real lengths rather than assuming.
- **Contradictory sources.** If two documents disagree, say which wins ("newer policy overrides older").
- **Instruction leakage from data.** Retrieved text can contain imperative sentences ("ignore previous instructions"). Delimiting helps, but see the prompt injection lesson for real defences.

## Context engineering for long-running assistants and agents

Context engineering matters most when context accumulates: multi-turn chats and agents that call tools dozens of times. Current platforms offer tools for this, and it helps to know the vocabulary:

- **Compaction:** summarising older turns into a compact state so the conversation can continue. Some APIs now offer server-side compaction; you can also do it yourself with a summarisation call.
- **Context editing or clearing:** removing stale tool results or old reasoning blocks that no longer matter, while keeping the decisions derived from them.
- **External memory:** writing notes, plans and facts to a file or store the agent can read back later, instead of keeping everything in the window.
- **Just-in-time retrieval:** giving the model lightweight references (file paths, IDs, search tools) and letting it load details only when needed, rather than pre-loading everything.

The design question is always the same: which tokens earn their place in this particular call?

## Hands-on: a context budget report

Measure before you optimise. This script counts tokens per component with the Claude token-counting endpoint, so you can see what dominates.

```python
import os
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-opus-5")

components = {
    "system": open("prompts/system.md", encoding="utf-8").read(),
    "retrieved_docs": open("context/retrieved.md", encoding="utf-8").read(),
    "history_summary": open("context/history.md", encoding="utf-8").read(),
    "question": "Can I return a sale item after 20 days?",
}

def count(text: str) -> int:
    result = client.messages.count_tokens(
        model=MODEL, messages=[{"role": "user", "content": text}]
    )
    return result.input_tokens

report = {name: count(text) for name, text in components.items()}
total = sum(report.values())
for name, n in sorted(report.items(), key=lambda kv: -kv[1]):
    print(f"{name:16} {n:7} tokens  {100 * n / total:5.1f}%")
```

Counts include a small per-message overhead, so treat them as close estimates. Token counts differ between model families because tokenisers differ; always count with the model you will use.

## Worked example: trimming a support bot

A support team finds that 70% of each request (illustrative figure) is old tool output: full order records returned by a lookup tool on earlier turns. They change the tool to return only the fields the model needs, replace earlier tool results with a one-line summary after each answer, and move stable policy text into a cached prefix. Answers get faster and cheaper, and accuracy on their eval set holds. The lesson: the biggest context wins are usually structural, not wording tweaks.

## Checklist before every production prompt

- Is every component needed for this call?
- Is anything stale, duplicated or contradictory?
- Are long documents above the question, labelled with source and date?
- Are stable parts first, so they can be cached?
- Do you know the token count of each part?

## Going further

Measure context, don't guess. Log token counts per component (system, history, retrieval, tools) for real traffic. You will often find that one component, such as verbose tool results, dominates cost while adding little. Trimming it is frequently the single cheapest quality and latency win available.

## Video lecture: Context engineering: what the model sees, and in what order

Lecture coming soon · 12 chapters · about 8 minutes. Read the full transcript below.

1. Context engineering
2. The core idea
3. Include, exclude, compress
4. Order and label
5. Long-context strategies
6. Worked example: the board pack
7. Context in long runs
8. Measure the context budget
9. Example 1: the dentist appointment
10. Example 2: a support bot's context (illustrative)
11. Failure modes
12. Recap

## Lecture transcript

### Context engineering

A strategy team pasted a one-hundred-and-twenty-page board pack into a model and asked: what are the risks? They got a generic list that could have described any company. The model was not the problem. What it was shown was. In this lecture you will learn context engineering: the discipline of curating everything the model sees, from instructions and documents to history, tools and tool results. You will learn to include, exclude and compress, how ordering and labelling change answers, the long-context strategies that work, and how to manage context in long-running agents.

### The core idea

As applications grew from single prompts into assistants and agents, practitioners started talking about context engineering. The context window holds the system prompt, retrieved documents, conversation history, tool definitions, tool results, examples and the user's request. The model can only reason about what it sees, and it reasons less reliably when what it sees is noisy, contradictory or bloated. So here is the core idea to remember: context is a scarce, expensive resource with diminishing returns. The goal is the smallest set of high-signal information that makes the correct answer likely.

### Include, exclude, compress

For every candidate piece of context, ask three questions. Does the task need it? A pricing question needs the price list, not the company history. Is it trustworthy and current? An outdated policy is worse than none, because the model will cite it confidently. And can it be compressed without losing signal? A forty-message chat can become a structured summary of decisions and open questions. People often forget to include the audience, today's date, the definition of done, and domain terms the model might misread, such as what active customer means in your company.

### Order and label

Position and structure matter. For long inputs, put the documents first and your instructions and question at the end; that tends to improve answer quality. Label everything with clear delimiters so the model can tell instructions from data. XML-style tags work well across many models: a documents tag containing each document with its source and content, then an instructions tag, then the question. And add metadata like source names, dates and authority levels, official policy versus forum post, so the model can weigh conflicting information. If two sources may disagree, state which wins, for example newer policy overrides older.

### Long-context strategies

When material is genuinely large, you have options, roughly from cheapest to most involved. Quote-then-answer: ask the model to extract the relevant quotes into a tagged section first, then answer from them. Map-reduce: process each chunk with the same prompt, then combine partial results. Retrieval: search for the most relevant passages and send only those. And structured memory: for long conversations, keep a running notes object of decisions, facts and open items, and pass that rather than the full transcript.

### Worked example: the board pack

Back to the board pack. The improved design removed appendices of raw tables the question did not need, added metadata saying which section was the finance report and which was the marketing update, and defined risk for this board: financial, regulatory and operational, with a materiality threshold. It asked for quotes with page references before the synthesis, and put the question and format at the end. The answer became shorter, specific and auditable. The model was given a sharper problem, not a bigger one.

### Context in long runs

Context engineering matters most when context accumulates, in long chats and in agents that call tools dozens of times. Learn the vocabulary. Compaction summarises older turns so work can continue; some APIs now offer it server-side. Context editing clears stale tool results or old reasoning blocks that no longer matter. External memory writes notes and plans to a file the agent reads back later. And just-in-time retrieval gives the model lightweight references like file paths or IDs, and lets it load details only when needed.

### Measure the context budget

Measure, do not guess. The lesson's hands-on script counts tokens for each component, system, retrieved documents, history summary and question, using the provider's token-counting endpoint, and prints the share of each. Count with the model you will use, because tokenisers differ between model families. Teams often discover that one component, typically verbose tool results or stale history, dominates cost while adding little. One support team found most of each request was old order records from earlier tool calls. Returning only needed fields and summarising old results made answers faster and cheaper with no loss on their eval set.

### Example 1: the dentist appointment

A simple worked example first. You ask a model: when is my next dentist appointment? It cannot know. Now add context: today's date, the clinic's confirmation email, and the note that you rescheduled once. The model answers correctly, and if you also pasted three old promotional emails from the clinic, it might have been confused by an outdated date in one of them. That is context engineering in miniature. Add what is needed, today's date and the latest confirmation. Exclude what misleads, the old promotions. And state precedence: if dates conflict, the most recent confirmation wins.

### Example 2: a support bot's context (illustrative)

Now a business scenario with illustrative numbers. An e-commerce brand's support bot sends about twenty-eight thousand tokens with every request: a full product catalogue, the whole returns policy history, the last forty messages, and verbose order records from tool calls. Answers are slow and sometimes cite old policies. The team measures each component. The catalogue is seventy percent of tokens but relevant to one question in ten. So they replace it with a product search tool. They keep only the current returns policy, with its effective date. They compress history into a short summary of the order, the issue and what has been tried. And they trim tool results to the five fields needed. Requests drop to around six thousand tokens. Latency and cost fall, and on their evaluation set accuracy rises slightly, because the model is no longer distracted by stale policy. All numbers illustrative.

### Failure modes

Watch for three failure modes. Context rot: quality can degrade as context grows, especially for details buried in the middle of very long inputs, so test at your real lengths. Contradictory sources without a precedence rule. And instruction leakage from data, where retrieved text contains imperatives like ignore previous instructions. Delimiting helps, but the prompt injection lesson covers real defences.

### Recap

To recap. Context engineering curates everything in the window. Aim for the smallest high-signal set. Put long material first and the question last, label everything with source and date, and state precedence rules. Use quote-then-answer, map-reduce, retrieval or structured memory for large inputs, and compaction, clearing and external notes for long runs. Try this now: pick one long-document task, rebuild it with labelled documents on top, a quote-extraction step and the question at the end, and compare accuracy on five questions you know the answers to. Next: designing few-shot examples.

## Key takeaways

- Context engineering curates everything in the window: instructions, documents, history, tools and results.
- Aim for the smallest set of high-signal context; more tokens add cost, latency and distraction.
- Put long material first, instructions and question last, and label each part with clear delimiters.
- For large inputs use quote-then-answer, map-reduce, retrieval or structured memory instead of dumping everything.

## Try it

Pick one long-document task you do. Rebuild the prompt with labelled documents at the top, a quote-extraction step, and the question at the end. Compare accuracy on five questions you know the answers to.

- [Previous: Designing system prompts and roles](https://optimizeall.com/learn/advanced-prompt-engineering/designing-system-prompts)
- [Next: Few-shot and example design](https://optimizeall.com/learn/advanced-prompt-engineering/few-shot-example-design)
- [All lessons of Advanced Prompt Engineering](https://optimizeall.com/learn/advanced-prompt-engineering)
