Building Production AI AgentsMemory, retrieval and durable state · Lesson 7 of 18

Context engineering and agent memory

Article · 16 min · 9 min lecture

Video lecture

Context engineering and agent memory

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Context engineering and memory

  • The window is working memory
  • Five memory types
  • Write paths, privacy and testing

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The context window is working memory

Everything an agent "knows" during a step is in its context window: system prompt, tool definitions, conversation history, tool results and any retrieved documents. Windows have grown large (hundreds of thousands to a million tokens on some current models), but bigger is not free: every token costs money on every call, and models attend less reliably to details buried in very long contexts. Context engineering is the discipline of putting the right tokens in the window at each step.

A taxonomy of agent memory

Memory typeWhat it holdsWhere it livesExample
Working (short-term)Current task, recent turns, latest tool resultsContext window"User wants Q3 report for UAE only"
EpisodicRecords of past sessions and outcomesDatabase / log store"Last month's report used the new attribution model"
SemanticFacts and knowledgeVector index, knowledge base, CRMProduct catalog, policies
ProceduralHow to do thingsPrompts, skills files, tool code"Monthly report checklist v3"
Profile / preferencesStable facts about a user or orgKey-value store"Prefers British spelling; fiscal year starts April"

Design question for each: who writes it (the agent, a human, a pipeline), who reads it, when does it expire, and how is it corrected.

Keeping working memory lean

  1. Concise tool outputs (lesson 4) are the biggest lever.
  2. Trim or clear stale tool results: once a search result has been used, the raw payload rarely matters. Anthropic offers context editing that clears old tool results automatically; you can also do it yourself by replacing old results with a one-line note.
  3. Compaction / summarization: when history nears a threshold, replace older turns with a summary. Anthropic offers server-side compaction in beta; frameworks such as the OpenAI Agents SDK and LangGraph provide session and summarization utilities. Test summaries: a bad summary silently drops a constraint the user gave in turn two.
  4. Structured state outside the prompt: keep facts like "approved budget = PKR 2.5m" in a state object and render them into the prompt each turn, rather than hoping the model remembers.
  5. Stable prefix for caching: put unchanging content (system prompt, tool list) first and volatile content last so prompt caching works (module 6).

Long-term memory: write paths matter most

Retrieval gets attention, but the hard part is deciding what to write. Good practice:

  • Write distilled facts, not transcripts: "Client ACME prefers Tuesday calls" rather than a 40-turn log.
  • Attach provenance (source, date, who confirmed) so stale or wrong memories can be traced and corrected.
  • Let users see and delete memories about them; this is both good UX and, in many jurisdictions (GDPR, UK GDPR, KSA's PDPL, the UAE's PDPL), a legal expectation for personal data.
  • Scope memory: per user, per team, per org. Never let one customer's memory leak into another's context.

Tools that implement memory

  • Anthropic's memory tool (a client-side tool type) lets Claude create, read and update files in a memory directory that you host; you control storage and access.
  • Managed platforms offer hosted memory stores for long-running agents.
  • Many teams simply expose remember(fact, scope) and recall(query, scope) tools backed by Postgres + pgvector or another vector store.

Worked example: an account-manager assistant for a Lahore agency

The assistant supports account managers across 40 clients.

  • Profile memory per client: brand voice, banned words, approval contacts, time zone.
  • Episodic memory: dated notes such as "Sept campaign paused after CPM spike; client asked for weekly updates".
  • Working memory: current request plus the three most relevant notes (retrieved by client ID and recency, then semantic similarity).
  • Write policy: the agent proposes new memories at the end of a session; the account manager approves them with one click.

Outcome: the assistant stops asking the same questions, and wrong memories are rare because humans approve writes.

Hands-on: a scoped memory store with provenance

import sqlite3, time, json

db = sqlite3.connect("memory.db")
db.execute("""CREATE TABLE IF NOT EXISTS memories(
  id INTEGER PRIMARY KEY, scope TEXT NOT NULL, fact TEXT NOT NULL,
  source TEXT NOT NULL, created REAL NOT NULL, approved INTEGER DEFAULT 0)""")

def propose_memory(scope: str, fact: str, source: str) -> int:
    cur = db.execute("INSERT INTO memories(scope, fact, source, created) VALUES (?,?,?,?)",
                     (scope, fact.strip()[:500], source, time.time()))
    db.commit()
    return cur.lastrowid

def approve(memory_id: int) -> None:
    db.execute("UPDATE memories SET approved=1 WHERE id=?", (memory_id,))
    db.commit()

def recall(scope: str, limit: int = 10) -> list[dict]:
    rows = db.execute("""SELECT fact, source, created FROM memories
                         WHERE scope=? AND approved=1 ORDER BY created DESC LIMIT ?""",
                      (scope, limit)).fetchall()
    return [{"fact": f, "source": s, "age_days": round((time.time() - c) / 86400)} for f, s, c in rows]

def render_memory_block(scope: str) -> str:
    items = recall(scope)
    lines = [f"- {m['fact']} (source: {m['source']}, {m['age_days']}d old)" for m in items]
    return "Known facts for this client (verify if older than 90 days):\n" + "\n".join(lines)

Expose propose_memory as a tool; keep approve for humans. Add semantic search later if recency plus scope is not enough.

Pitfalls

  • Stuffing entire histories into context "just in case".
  • Memories without dates or sources, which can never be audited.
  • Cross-tenant leakage through a shared vector index without filters.
  • Summaries that drop constraints; test compaction with adversarial long conversations.

Measuring success

Track repeated-question rate, memory precision (share of recalled memories judged relevant), stale-memory incidents, tokens per step before and after trimming, and task success on long conversations.

Key takeaways

  • Context engineering means putting the right tokens, not the most tokens, in the window each step.
  • Separate working, episodic, semantic, procedural and profile memory, each with clear owners and expiry.
  • Keep working memory lean with concise tool outputs, clearing stale results and tested compaction.
  • The write path matters most: store distilled, dated facts with provenance, scoped per tenant.
  • Let people see and delete memories about them, which is good UX and often a legal requirement.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent forgets a budget limit the user stated early in a long conversation after compaction. What is the most robust fix?
  2. Which memory write is best?
  3. What risk does a shared vector index without tenant filters create?

Put it into practice

List the five memory types for an assistant you want to build. For each, define who writes it, who reads it, when it expires and how a user can correct it.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.