Latest AI Techniques: RAG, Tool Use, Agents & MCPAgent loops, multi-agent patterns and memory · Lesson 9 of 20
Agent loops: plan, act, observe
Video lecture
Agent loops: plan, act, observe
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Agent loops
Agent might be the most overused word in AI marketing right now. Everything from a chatbot to a scheduled email gets called one. In this lesson you'll get a precise, practical definition, see the loop that every agent runs, learn the stop conditions that separate agents that work from agents that embarrass you, and walk through a hands-on loop with approvals and loop detection built in.
0:29 Analogy: train vs taxi
Why does the definition matter? Because agents cost more, run slower and fail in more interesting ways than workflows. Here's an analogy. A workflow is a train on rails: it goes exactly where the track goes, reliably. An agent is a taxi driver: you give a destination, and they choose the route, adapting to traffic. Taxis are wonderful when you don't know the route in advance. But you'd never use one for a journey the train already runs perfectly.
1:03 Agent vs workflow
Here's the definition. An agent is a system where the model directs its own process. It decides which tool to use next, looks at the result, and decides again, in a loop, until the goal is met or it hits a limit. A workflow is different: the steps are fixed in code and the model fills in individual steps. Both are valuable. Workflows are predictable and cheap. Agents shine when the path genuinely can't be known in advance, like research, debugging or open-ended customer requests.
1:40 Plan → act → observe
The loop is plan, act, observe. The model states a short plan, takes an action like a search or a database lookup, observes what comes back, and updates the plan. Asking for an explicit plan at the start, and a quick revision as observations arrive, helps the agent stay oriented on longer tasks. The model's reasoning is only as good as the feedback it gets, so every tool result should say clearly whether the step worked, like tests passed, zero rows returned, or validation failed on the date field.
2:19 Stop conditions
Now the part people skip: stop conditions. Every agent needs a maximum number of steps or tool calls; a budget in tokens, money or time; success criteria it can check, ideally in code, like the report contains all five sections with sources; and escalation rules, such as if you can't find the order after two searches, hand off to a human. Without these, agents loop, repeat failing actions and quietly run up bills.
2:51 Simple example: 'Has my order shipped?'
A simple example before the bigger one. The goal: find out whether a customer's order has shipped, and if not, why. The agent plans: look up the order, then check the warehouse if needed. It calls get order and sees status pending. It calls check warehouse and sees out of stock, restock Friday. It checks its definition of done, status and reason found, and replies to the customer's agent with both. Two tools, two steps, and a path that depended on what the first step revealed.
3:28 Worked example: pricing research
Let's watch one. The goal: a one-page comparison of three competitors' pricing pages. The agent plans: find the pages, extract tiers, compare. It opens competitor A: three tiers, easy. Competitor B hides pricing behind contact sales. A good agent records not public instead of guessing. It extracts C, checks its definition of done, all three covered with sources and gaps noted, and delivers with citations. A fixed workflow would have broken on B. The agent adapted, but within guardrails.
4:02 Business example (illustrative)
A deeper business example, illustrative. A Dubai product team ran the pricing research agent weekly across eight competitors. Each run took about twelve minutes and cost less than a coffee, versus half a day of an analyst's time. Twice in the first month it recorded pricing not public rather than guessing, which the analyst then checked by phone. The analyst now reviews a finished draft for twenty minutes instead of compiling it from scratch.
4:34 Failure modes
Common failure modes are worth memorising. Looping, repeating the same failed action, which you detect by spotting identical calls. Premature completion, declaring victory early, which you counter with explicit checks. Goal drift, getting lost in a sub-task, so restate the goal periodically. Compounding errors, where an early wrong assumption poisons later steps. And unsafe actions triggered by manipulated content, which you prevent with approval gates and by treating retrieved content as untrusted.
5:05 Autonomy is a dial
Treat autonomy as a dial, not a switch. Level one, suggest only. Level two, act with approval for each consequential step. Level three, act autonomously inside a sandbox or budget with review afterwards. Level four, act autonomously in production. Most business deployments today sit at levels two and three, and you move up only as evaluation and monitoring earn it. The lesson's hands-on code gives you a real loop with the Claude API: an approval gate for sending emails or refunds, detection of repeated identical calls, a token budget, and error results returned to the model instead of crashing.
5:48 Common mistakes
Common mistakes when building agent loops. No step limit, so a confused agent keeps going. Tools that return vague errors, so the agent can't recover. Asking for a plan but never letting it be revised. Treating a model's claim of done as proof, instead of checking in code. And jumping straight to high autonomy because the demo looked good. Every one of these shows up in real incident reports.
6:18 How you'll know it's working
How will you know your agent loop is healthy? Success rate against a checkable definition of done, measured over repeated runs. Median steps and tokens per task, which should be stable, not creeping upward. The number of loop-detection or step-limit stops, which should be rare. And the escalation rate: some escalation is healthy, because it means the agent knows when it's stuck.
6:45 Watch me do it: run_agent()
Watch me do it. I open run agent. First, the system prompt asks for a three to five line plan and to stop when done or explain the blocker. Next, inside the loop, I call the API, add input and output tokens to a running total, and append the assistant content. If the stop reason isn't tool use, I return done with the answer, steps and tokens. If the token budget is exceeded, I stop with reason budget. For each tool call, I build a key from the tool name and sorted JSON input. If I've seen that exact key, I return an error saying change approach. If the tool needs approval and the approver says no, I return a declined error. Otherwise I run the handler and catch any exception as an error result. I test it with a fake send email tool and an approver that always says no. The trace shows the agent drafting instead, and the run ends cleanly in three steps.
7:57 Recap
To recap: an agent lets the model choose its next step in a loop; a workflow fixes the steps. The loop is plan, act, observe, with informative feedback, and it always needs step limits, budgets, a checkable definition of done and escalation rules. Your next step is to design one agent on paper: goal, definition of done, tools, step limit, budget and two escalation rules. Then compare it with a simple workflow and ask honestly which one you need. For a deep build, see the Building Production AI Agents course.
8:36 Try this now (15 minutes)
Try this now. Pick one open-ended task from your work where the steps really do depend on what you discover, like researching a prospect or investigating a customer complaint. On one page, write the goal, a checkable definition of done, the tools it needs, a step limit, a budget and two escalation rules. Then ask honestly: could a fixed workflow do this? If yes, build that instead.
What makes something an agent
A useful working definition: an agent is a system where a model directs its own process, choosing which tools to use and what to do next, in a loop, until a goal is met or a stop condition is reached. Contrast this with a workflow, where the sequence of steps is fixed in code and the model fills in individual steps.
Both are valuable. Workflows are predictable and cheap. Agents are flexible when the path cannot be known in advance, such as research, debugging, or handling open-ended customer requests.
The core loop
def run_agent(goal, tools, max_steps=15):
messages = [system_prompt(), user(goal)]
for step in range(max_steps):
reply = model.generate(messages, tools=tools)
messages.append(reply)
if reply.is_final_answer():
return reply.text
for call in reply.tool_calls:
if needs_approval(call) and not human_approves(call):
result = "User declined this action."
else:
result = execute_safely(call) # validate, authorise, run
messages.append(tool_result(call.id, result))
return "Stopped: step limit reached. Summary of progress: ..."The pattern is often described as plan, act, observe: the model reasons about what to do, calls a tool, observes the result, and repeats. Many teams find it helpful to ask the model to state a brief plan at the start and revise it as observations arrive.
Essential stop conditions
Agents without limits can loop, repeat failing actions or run up costs. Always set:
- a maximum number of steps or tool calls
- a budget (tokens, money or time)
- success criteria the model can check ("done when the report includes all five sections with sources")
- escalation rules ("if you cannot find the order after two searches, hand off to a human")
The agent-computer interface
An agent is only as good as its tools, its instructions and the information it sees each turn. Practical guidance:
- Clear goal and definition of done in the system prompt.
- Tools designed for the task (previous module), with results that are concise and informative.
- Environment feedback: make results tell the agent whether the step worked (test results, validation errors, "0 rows returned").
- Progress tracking: for long tasks, have the agent maintain a short plan or checklist that it updates, so it stays oriented.
Worked example: a competitive research agent
Goal: "Produce a one-page comparison of three competitors' pricing pages for our product team."
- Plan: list competitors, locate pricing pages, extract tiers, compare.
- Act: web search, open page. Observe: page A has three tiers.
- Act: open page B. Observe: pricing hidden behind "contact sales". The agent records "not public" instead of guessing.
- Act: open page C; extract.
- Check against definition of done: all three covered, sources linked, gaps noted.
- Final answer with citations and a note of what could not be verified.
A fixed workflow would have struggled with B's missing pricing; the agent adapted. But notice the guardrails: it records unknowns rather than inventing, and it cites sources.
Common agent failure modes
- Looping: repeating the same failed action. Detect repeated identical calls and intervene.
- Premature completion: declaring success without meeting criteria. Add explicit checks, ideally verifiable in code.
- Goal drift: getting absorbed in a sub-task. Periodically restate the goal and plan.
- Compounding errors: an early wrong assumption poisons later steps. Encourage verification of key facts before acting on them.
- Unsafe actions: taking consequential actions based on manipulated content. Gate actions and treat retrieved content as untrusted.
Autonomy is a dial, not a switch
Think of autonomy levels:
- Suggest only (human does everything).
- Act with approval for each consequential step.
- Act autonomously within a sandbox or budget, with review after.
- Act autonomously in production.
Most business deployments today sit at levels 2 and 3. Move up the dial only as evaluations and monitoring build justified confidence.
Hands-on: a bounded agent loop with approvals and loop detection
This is the pseudo-code from above made real with the Claude API. It adds three things production agents need: an approval gate for consequential tools, detection of repeated identical calls, and a budget measured in tokens.
import json, os
import anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
NEEDS_APPROVAL = {"send_email", "issue_refund"}
TOKEN_BUDGET = 60_000
def run_agent(goal, tools, handlers, max_steps=15, approve=lambda name, args: False):
system = ("Work step by step towards the goal. Start with a 3-5 line plan and update it as you learn. "
"Stop when the definition of done is met, or explain what blocked you.")
messages = [{"role": "user", "content": goal}]
seen, used = set(), 0
for step in range(max_steps):
resp = client.messages.create(model=MODEL, max_tokens=4000, system=system,
tools=tools, messages=messages)
used += resp.usage.input_tokens + resp.usage.output_tokens
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use":
return {"status": "done", "answer": "".join(b.text for b in resp.content if b.type == "text"),
"steps": step + 1, "tokens": used}
if used > TOKEN_BUDGET:
return {"status": "stopped", "reason": "budget", "steps": step + 1, "tokens": used}
results = []
for call in (b for b in resp.content if b.type == "tool_use"):
key = (call.name, json.dumps(call.input, sort_keys=True))
if key in seen:
out, err = "You already made this exact call. Change approach or explain the blocker.", True
elif call.name in NEEDS_APPROVAL and not approve(call.name, call.input):
out, err = "A human declined this action. Do not retry it; propose an alternative.", True
else:
try:
out, err = json.dumps(handlers[call.name](**call.input)), False
except Exception as e: # return errors to the model instead of crashing the loop
out, err = f"{type(e).__name__}: {e}", True
seen.add(key)
results.append({"type": "tool_result", "tool_use_id": call.id, "content": out, "is_error": err})
messages.append({"role": "user", "content": results})
return {"status": "stopped", "reason": "step_limit", "tokens": used}Log every run's goal, steps, tool calls, approvals and final status. Those logs become the raw material for the agent evaluation you will build in module 6.
Managed harnesses
You do not always need to own this loop. Current options include SDK helpers that run the loop for tools you define (for example the Anthropic SDK's tool runner), full agent harnesses you host yourself (such as the Claude Agent SDK or the OpenAI Agents SDK), and hosted agent platforms where the provider runs both the loop and a sandbox. Lesson "Agent frameworks and SDKs" compares them. Whatever you choose, the ideas here, stop conditions, approvals, informative feedback and autonomy as a dial, still apply.
Going further
Long-running agents face context limits: the transcript of tool calls grows until it no longer fits or becomes noisy. Techniques include summarising or compacting older turns, storing notes in external files or memory tools, and splitting work across sub-agents with fresh context (next lesson). Provider platforms increasingly offer built-in support for these; the underlying ideas are stable even as features change.
Key takeaways
- An agent lets the model choose its next steps in a loop; a workflow fixes the steps in code.
- The loop is plan, act, observe, with stop conditions: step limits, budgets, success criteria and escalation.
- Agent quality depends on clear goals, task-fit tools, informative feedback and progress tracking.
- Treat autonomy as a dial and increase it only as evaluation and monitoring justify.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Design an agent for one open-ended task. Write its goal, definition of done, tools, step limit, budget and two escalation rules.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.