Latest AI Techniques: RAG, Tool Use, Agents & MCPAgent loops, multi-agent patterns and memory · Lesson 9 of 20

Agent loops: plan, act, observe

Video lesson · 13 min · 9 min lecture

Video lecture

Agent loops: plan, act, observe

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Agent loops

  • A precise definition
  • Plan, act, observe
  • Stop conditions and autonomy levels

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What makes something an agent

A useful working definition: an agent is a system where a model directs its own process, choosing which tools to use and what to do next, in a loop, until a goal is met or a stop condition is reached. Contrast this with a workflow, where the sequence of steps is fixed in code and the model fills in individual steps.

Both are valuable. Workflows are predictable and cheap. Agents are flexible when the path cannot be known in advance, such as research, debugging, or handling open-ended customer requests.

The core loop

def run_agent(goal, tools, max_steps=15):
    messages = [system_prompt(), user(goal)]
    for step in range(max_steps):
        reply = model.generate(messages, tools=tools)
        messages.append(reply)
        if reply.is_final_answer():
            return reply.text
        for call in reply.tool_calls:
            if needs_approval(call) and not human_approves(call):
                result = "User declined this action."
            else:
                result = execute_safely(call)       # validate, authorise, run
            messages.append(tool_result(call.id, result))
    return "Stopped: step limit reached. Summary of progress: ..."

The pattern is often described as plan, act, observe: the model reasons about what to do, calls a tool, observes the result, and repeats. Many teams find it helpful to ask the model to state a brief plan at the start and revise it as observations arrive.

Essential stop conditions

Agents without limits can loop, repeat failing actions or run up costs. Always set:

  • a maximum number of steps or tool calls
  • a budget (tokens, money or time)
  • success criteria the model can check ("done when the report includes all five sections with sources")
  • escalation rules ("if you cannot find the order after two searches, hand off to a human")

The agent-computer interface

An agent is only as good as its tools, its instructions and the information it sees each turn. Practical guidance:

  • Clear goal and definition of done in the system prompt.
  • Tools designed for the task (previous module), with results that are concise and informative.
  • Environment feedback: make results tell the agent whether the step worked (test results, validation errors, "0 rows returned").
  • Progress tracking: for long tasks, have the agent maintain a short plan or checklist that it updates, so it stays oriented.

Worked example: a competitive research agent

Goal: "Produce a one-page comparison of three competitors' pricing pages for our product team."

  1. Plan: list competitors, locate pricing pages, extract tiers, compare.
  2. Act: web search, open page. Observe: page A has three tiers.
  3. Act: open page B. Observe: pricing hidden behind "contact sales". The agent records "not public" instead of guessing.
  4. Act: open page C; extract.
  5. Check against definition of done: all three covered, sources linked, gaps noted.
  6. Final answer with citations and a note of what could not be verified.

A fixed workflow would have struggled with B's missing pricing; the agent adapted. But notice the guardrails: it records unknowns rather than inventing, and it cites sources.

Common agent failure modes

  • Looping: repeating the same failed action. Detect repeated identical calls and intervene.
  • Premature completion: declaring success without meeting criteria. Add explicit checks, ideally verifiable in code.
  • Goal drift: getting absorbed in a sub-task. Periodically restate the goal and plan.
  • Compounding errors: an early wrong assumption poisons later steps. Encourage verification of key facts before acting on them.
  • Unsafe actions: taking consequential actions based on manipulated content. Gate actions and treat retrieved content as untrusted.

Autonomy is a dial, not a switch

Think of autonomy levels:

  1. Suggest only (human does everything).
  2. Act with approval for each consequential step.
  3. Act autonomously within a sandbox or budget, with review after.
  4. Act autonomously in production.

Most business deployments today sit at levels 2 and 3. Move up the dial only as evaluations and monitoring build justified confidence.

Hands-on: a bounded agent loop with approvals and loop detection

This is the pseudo-code from above made real with the Claude API. It adds three things production agents need: an approval gate for consequential tools, detection of repeated identical calls, and a budget measured in tokens.

import json, os
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")
NEEDS_APPROVAL = {"send_email", "issue_refund"}
TOKEN_BUDGET = 60_000

def run_agent(goal, tools, handlers, max_steps=15, approve=lambda name, args: False):
    system = ("Work step by step towards the goal. Start with a 3-5 line plan and update it as you learn. "
              "Stop when the definition of done is met, or explain what blocked you.")
    messages = [{"role": "user", "content": goal}]
    seen, used = set(), 0
    for step in range(max_steps):
        resp = client.messages.create(model=MODEL, max_tokens=4000, system=system,
                                      tools=tools, messages=messages)
        used += resp.usage.input_tokens + resp.usage.output_tokens
        messages.append({"role": "assistant", "content": resp.content})
        if resp.stop_reason != "tool_use":
            return {"status": "done", "answer": "".join(b.text for b in resp.content if b.type == "text"),
                    "steps": step + 1, "tokens": used}
        if used > TOKEN_BUDGET:
            return {"status": "stopped", "reason": "budget", "steps": step + 1, "tokens": used}
        results = []
        for call in (b for b in resp.content if b.type == "tool_use"):
            key = (call.name, json.dumps(call.input, sort_keys=True))
            if key in seen:
                out, err = "You already made this exact call. Change approach or explain the blocker.", True
            elif call.name in NEEDS_APPROVAL and not approve(call.name, call.input):
                out, err = "A human declined this action. Do not retry it; propose an alternative.", True
            else:
                try:
                    out, err = json.dumps(handlers[call.name](**call.input)), False
                except Exception as e:  # return errors to the model instead of crashing the loop
                    out, err = f"{type(e).__name__}: {e}", True
            seen.add(key)
            results.append({"type": "tool_result", "tool_use_id": call.id, "content": out, "is_error": err})
        messages.append({"role": "user", "content": results})
    return {"status": "stopped", "reason": "step_limit", "tokens": used}

Log every run's goal, steps, tool calls, approvals and final status. Those logs become the raw material for the agent evaluation you will build in module 6.

Managed harnesses

You do not always need to own this loop. Current options include SDK helpers that run the loop for tools you define (for example the Anthropic SDK's tool runner), full agent harnesses you host yourself (such as the Claude Agent SDK or the OpenAI Agents SDK), and hosted agent platforms where the provider runs both the loop and a sandbox. Lesson "Agent frameworks and SDKs" compares them. Whatever you choose, the ideas here, stop conditions, approvals, informative feedback and autonomy as a dial, still apply.

Going further

Long-running agents face context limits: the transcript of tool calls grows until it no longer fits or becomes noisy. Techniques include summarising or compacting older turns, storing notes in external files or memory tools, and splitting work across sub-agents with fresh context (next lesson). Provider platforms increasingly offer built-in support for these; the underlying ideas are stable even as features change.

Key takeaways

  • An agent lets the model choose its next steps in a loop; a workflow fixes the steps in code.
  • The loop is plan, act, observe, with stop conditions: step limits, budgets, success criteria and escalation.
  • Agent quality depends on clear goals, task-fit tools, informative feedback and progress tracking.
  • Treat autonomy as a dial and increase it only as evaluation and monitoring justify.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What distinguishes an agent from a fixed workflow?
  2. An agent keeps calling the same failing search with identical arguments. What safeguard helps?
  3. For most business deployments today, which autonomy level is typical?

Put it into practice

Design an agent for one open-ended task. Write its goal, definition of done, tools, step limit, budget and two escalation rules.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.