Building Production AI AgentsMemory, retrieval and durable state · Lesson 7 of 18
Context engineering and agent memory
Video lecture
Context engineering and agent memory
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Context engineering and memory
Have you ever had a colleague who forgets what you told them yesterday, and another who remembers everything but can't tell what matters? Agents can be both. In this lesson you'll learn context engineering, which is choosing the right tokens for each step, and how to design memory that is useful, accurate and safe.
0:23 Why it matters
Why does this matter? Because most long running agents fail not from lack of intelligence, but from a messy context. Picture your desk during a big project. If every document you've ever touched is piled on it, you can't find the one page that matters. If you file everything away, you keep walking to the cabinet. Good context engineering is a tidy desk: today's task in the middle, the few relevant notes beside it, everything else filed with a clear label. Agents need exactly that discipline.
1:00 The context window
Everything an agent knows during a step sits in its context window: the system prompt, tool definitions, the conversation so far, tool results and retrieved documents. Windows are large now, up to around a million tokens on some current models. But more isn't free. Every token costs money on every call, and models attend less reliably to details buried deep in a very long context. So the skill is curation, not stuffing.
1:31 Five memory types
There are five kinds of memory to design. Working memory is the current task and recent turns, living in the context. Episodic memory records past sessions and outcomes. Semantic memory holds facts, like your product catalog or policies. Procedural memory is how to do things: prompts, skills files and tool code. And profile memory holds stable preferences, like prefers British spelling or fiscal year starts in April. For each one, decide who writes it, who reads it, when it expires, and how it gets corrected.
2:08 Keep working memory lean
Keeping working memory lean is mostly about habits. Make tool outputs concise. Clear stale tool results once they've been used; Anthropic offers context editing that does this automatically, or you can replace old results with a one line note. When history gets long, compact it into a summary, but test those summaries, because a bad one quietly drops a constraint from turn two. Keep critical facts, like an approved budget, in a structured state object that you render into every prompt. And keep the unchanging parts at the front so prompt caching works.
2:48 Simple example: the forgotten currency
A simple example. A user tells a support agent in the first message: I'm on the Pro plan and I'm in Pakistan, so prices in rupees please. Forty messages later, after lots of troubleshooting, the conversation gets summarized to save space, and the summary forgets the currency. The agent quotes dollars. The fix isn't a smarter summarizer. It's storing plan equals Pro and currency equals PKR in a small state object that's rendered into every prompt. Critical facts go in state, not in the hope that a summary keeps them.
3:27 Write path rules
With long term memory, everyone obsesses over retrieval. But the hard part is deciding what to write. Write distilled facts, not transcripts. Client prefers Tuesday calls beats a forty turn log. Attach provenance, meaning the source, the date and who confirmed it, so wrong memories can be traced and fixed. Scope memory per user, per team or per organization, and never let one customer's memory leak into another's context. And let people see and delete what's remembered about them. That's good experience design, and under laws like GDPR, UK GDPR and the data protection laws in Saudi Arabia and the UAE, it's often expected for personal data.
4:14 Example: agency account assistant
Here's a worked example. A Lahore agency's account managers look after forty clients. Each client has profile memory: brand voice, banned words, approval contacts and time zone. Episodic notes record things like: September campaign paused after a cost spike, client wants weekly updates. For each request, the assistant pulls the three most relevant notes for that client. And at the end of a session, it proposes new memories, which the account manager approves with one click. The result: fewer repeated questions, and very few wrong memories.
4:51 Hands-on: scoped memory store
The hands on code builds a small memory store in SQLite. Memories have a scope, a fact, a source and a timestamp, and they start unapproved. The agent gets a propose memory tool. Humans keep the approve function. Recall returns only approved facts for one scope, with their age, and the prompt tells the model to verify anything older than ninety days. It's simple, but it already has the three things most memory systems lack: provenance, scoping and a human in the write path.
5:28 Pitfalls and metrics
Watch out for the classic mistakes: stuffing entire histories into context just in case, memories with no dates or sources, cross tenant leakage through a shared index without filters, and summaries that silently drop constraints. Measure repeated question rate, how often recalled memories are actually relevant, stale memory incidents, and tokens per step before and after trimming.
5:53 Auto-write or approve?
A common question: should an agent write to memory automatically, or should a human approve? It depends on the stakes. For low risk preferences, like prefers bullet points, automatic writes with an easy undo are fine. For facts that change decisions, like a client's budget, contract terms or who has approval authority, have the agent propose and a person confirm. And whatever you choose, show users what's remembered and let them edit it. Visible memory builds trust; invisible memory creates surprises.
6:28 Deeper: the Lahore agency memory rollout
Let's deepen the Lahore agency example. Forty clients, eight account managers, and before the change each manager repeated the same background questions in almost every AI session: what's this client's tone, who approves posts, which words are banned. After adding profile memory and approved episodic notes, the assistant pulled the three most relevant notes for each request. In the first two months, illustrative again, the number of repeated setup questions fell sharply, and the managers approved about eight out of ten proposed memories unchanged. The two in ten they rejected were telling: mostly one off comments that shouldn't be remembered, like a client being annoyed on a particular day. Human approval kept those out of the permanent record.
7:19 Watch me do it: the memory store
Watch me do it. Let's walk through the memory store. The table has an id, a scope, the fact, its source, a timestamp and an approved flag that starts at zero. Propose memory inserts a trimmed fact of at most five hundred characters with its source and time, and returns the new id. This is the only function I expose to the agent as a tool. Approve flips the flag, and only humans can call it, from the account manager's screen. Recall selects approved facts for one scope, newest first, and returns each with its source and age in days. Finally, render memory block turns them into a short prompt section, and the heading tells the model to verify anything older than ninety days. I'll run it: propose client prefers Tuesday calls for scope noor cosmetics, approve it, then render the block. There it is, with its source and age.
8:24 Try this now
Try this now. Take an assistant you use or want to build, and make a five row table, one row per memory type: working, episodic, semantic, procedural and profile. For each, write who writes it, who reads it, when it expires, and how a user can correct or delete it. Then pick the single most important fact your assistant keeps forgetting, and move it into structured state. That one change often fixes the most annoying bug users complain about.
8:58 Recap
To recap: curate the context window, design five memory types deliberately, keep working memory lean, and treat the write path as the hard part, with provenance, scoping and deletion. Your next step: list the five memory types for an assistant you want to build, and for each one define who writes it, who reads it, when it expires and how it's corrected.
The context window is working memory
Everything an agent "knows" during a step is in its context window: system prompt, tool definitions, conversation history, tool results and any retrieved documents. Windows have grown large (hundreds of thousands to a million tokens on some current models), but bigger is not free: every token costs money on every call, and models attend less reliably to details buried in very long contexts. Context engineering is the discipline of putting the right tokens in the window at each step.
A taxonomy of agent memory
| Memory type | What it holds | Where it lives | Example |
|---|---|---|---|
| Working (short-term) | Current task, recent turns, latest tool results | Context window | "User wants Q3 report for UAE only" |
| Episodic | Records of past sessions and outcomes | Database / log store | "Last month's report used the new attribution model" |
| Semantic | Facts and knowledge | Vector index, knowledge base, CRM | Product catalog, policies |
| Procedural | How to do things | Prompts, skills files, tool code | "Monthly report checklist v3" |
| Profile / preferences | Stable facts about a user or org | Key-value store | "Prefers British spelling; fiscal year starts April" |
Design question for each: who writes it (the agent, a human, a pipeline), who reads it, when does it expire, and how is it corrected.
Keeping working memory lean
- Concise tool outputs (lesson 4) are the biggest lever.
- Trim or clear stale tool results: once a search result has been used, the raw payload rarely matters. Anthropic offers context editing that clears old tool results automatically; you can also do it yourself by replacing old results with a one-line note.
- Compaction / summarization: when history nears a threshold, replace older turns with a summary. Anthropic offers server-side compaction in beta; frameworks such as the OpenAI Agents SDK and LangGraph provide session and summarization utilities. Test summaries: a bad summary silently drops a constraint the user gave in turn two.
- Structured state outside the prompt: keep facts like "approved budget = PKR 2.5m" in a state object and render them into the prompt each turn, rather than hoping the model remembers.
- Stable prefix for caching: put unchanging content (system prompt, tool list) first and volatile content last so prompt caching works (module 6).
Long-term memory: write paths matter most
Retrieval gets attention, but the hard part is deciding what to write. Good practice:
- Write distilled facts, not transcripts: "Client ACME prefers Tuesday calls" rather than a 40-turn log.
- Attach provenance (source, date, who confirmed) so stale or wrong memories can be traced and corrected.
- Let users see and delete memories about them; this is both good UX and, in many jurisdictions (GDPR, UK GDPR, KSA's PDPL, the UAE's PDPL), a legal expectation for personal data.
- Scope memory: per user, per team, per org. Never let one customer's memory leak into another's context.
Tools that implement memory
- Anthropic's memory tool (a client-side tool type) lets Claude create, read and update files in a memory directory that you host; you control storage and access.
- Managed platforms offer hosted memory stores for long-running agents.
- Many teams simply expose
remember(fact, scope)andrecall(query, scope)tools backed by Postgres + pgvector or another vector store.
Worked example: an account-manager assistant for a Lahore agency
The assistant supports account managers across 40 clients.
- Profile memory per client: brand voice, banned words, approval contacts, time zone.
- Episodic memory: dated notes such as "Sept campaign paused after CPM spike; client asked for weekly updates".
- Working memory: current request plus the three most relevant notes (retrieved by client ID and recency, then semantic similarity).
- Write policy: the agent proposes new memories at the end of a session; the account manager approves them with one click.
Outcome: the assistant stops asking the same questions, and wrong memories are rare because humans approve writes.
Hands-on: a scoped memory store with provenance
import sqlite3, time, json
db = sqlite3.connect("memory.db")
db.execute("""CREATE TABLE IF NOT EXISTS memories(
id INTEGER PRIMARY KEY, scope TEXT NOT NULL, fact TEXT NOT NULL,
source TEXT NOT NULL, created REAL NOT NULL, approved INTEGER DEFAULT 0)""")
def propose_memory(scope: str, fact: str, source: str) -> int:
cur = db.execute("INSERT INTO memories(scope, fact, source, created) VALUES (?,?,?,?)",
(scope, fact.strip()[:500], source, time.time()))
db.commit()
return cur.lastrowid
def approve(memory_id: int) -> None:
db.execute("UPDATE memories SET approved=1 WHERE id=?", (memory_id,))
db.commit()
def recall(scope: str, limit: int = 10) -> list[dict]:
rows = db.execute("""SELECT fact, source, created FROM memories
WHERE scope=? AND approved=1 ORDER BY created DESC LIMIT ?""",
(scope, limit)).fetchall()
return [{"fact": f, "source": s, "age_days": round((time.time() - c) / 86400)} for f, s, c in rows]
def render_memory_block(scope: str) -> str:
items = recall(scope)
lines = [f"- {m['fact']} (source: {m['source']}, {m['age_days']}d old)" for m in items]
return "Known facts for this client (verify if older than 90 days):\n" + "\n".join(lines)Expose propose_memory as a tool; keep approve for humans. Add semantic search later if recency plus scope is not enough.
Pitfalls
- Stuffing entire histories into context "just in case".
- Memories without dates or sources, which can never be audited.
- Cross-tenant leakage through a shared vector index without filters.
- Summaries that drop constraints; test compaction with adversarial long conversations.
Measuring success
Track repeated-question rate, memory precision (share of recalled memories judged relevant), stale-memory incidents, tokens per step before and after trimming, and task success on long conversations.
Key takeaways
- Context engineering means putting the right tokens, not the most tokens, in the window each step.
- Separate working, episodic, semantic, procedural and profile memory, each with clear owners and expiry.
- Keep working memory lean with concise tool outputs, clearing stale results and tested compaction.
- The write path matters most: store distilled, dated facts with provenance, scoped per tenant.
- Let people see and delete memories about them, which is good UX and often a legal requirement.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
List the five memory types for an assistant you want to build. For each, define who writes it, who reads it, when it expires and how a user can correct it.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.