Building Production AI AgentsAgent foundations and architectures · Lesson 3 of 18

Build the agent loop from scratch in Python

Article · 18 min · 9 min lecture

Video lecture

Build the agent loop from scratch in Python

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Build the agent loop from scratch

  • The anatomy of a tool call
  • A production-minded loop
  • Budgets, errors and logs

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why build it by hand first

Frameworks and SDK helpers are great, but if you have never written the loop yourself you will struggle to debug them. The raw loop is about 40 lines. Once you understand it, every framework is just conveniences around it.

The anatomy of a tool call (Claude Messages API)

  1. You send messages, a system prompt and a list of tools, each with a name, description and JSON Schema input_schema.
  2. The model replies. If it wants a tool, stop_reason is "tool_use" and the content contains one or more tool_use blocks with an id, name and input.
  3. You append the assistant's full content (not just its text) to messages, run each tool, then append one user message containing a tool_result block per call, each referencing the matching tool_use_id.
  4. Call the API again. Repeat until stop_reason is "end_turn" (or something else you must handle).

Other stop reasons to handle: "max_tokens" (output was cut off; raise the limit or ask it to continue), "pause_turn" (a long server-side tool turn paused; send the conversation back to continue), and "refusal" on some newer models (the model declined; do not read tool calls from it). OpenAI's Responses API and Google's Gemini API follow the same idea with different field names, which the platform-integration course compares side by side.

Hands-on: a minimal but production-minded loop

import os, json, time, logging
import anthropic

logging.basicConfig(level=logging.INFO)
log = logging.getLogger("agent")
client = anthropic.Anthropic()  # ANTHROPIC_API_KEY from env; SDK retries 429/5xx by default
MODEL = os.environ.get("AGENT_MODEL", "claude-sonnet-5")  # verify current IDs in the docs

TOOLS = [
    {
        "name": "search_crm",
        "description": "Search CRM contacts by company name. Returns up to 5 matches with id, name, role, last_contacted.",
        "input_schema": {
            "type": "object",
            "properties": {"company": {"type": "string", "description": "Company name, e.g. 'Acme Ltd'"}},
            "required": ["company"],
            "additionalProperties": False,
        },
    },
    {
        "name": "get_deals",
        "description": "List open deals for a CRM contact id. Returns stage, value and next step.",
        "input_schema": {
            "type": "object",
            "properties": {"contact_id": {"type": "string"}},
            "required": ["contact_id"],
            "additionalProperties": False,
        },
    },
]

def search_crm(company: str) -> list[dict]:        # replace with your real CRM client
    return [{"id": "c_101", "name": "Sara Khan", "role": "Head of Growth", "last_contacted": "2026-08-30"}]

def get_deals(contact_id: str) -> list[dict]:
    return [{"stage": "proposal", "value_gbp": 12000, "next_step": "pricing call"}]

REGISTRY = {"search_crm": search_crm, "get_deals": get_deals}

def run_tool(name: str, args: dict) -> tuple[str, bool]:
    fn = REGISTRY.get(name)
    if fn is None:
        return f"Unknown tool {name}", True
    try:
        return json.dumps(fn(**args))[:8000], False   # cap tool output size
    except Exception as exc:                          # return errors to the model, don't crash
        return f"Tool error: {type(exc).__name__}: {exc}", True

def run_agent(goal: str, max_steps: int = 10) -> str:
    messages = [{"role": "user", "content": goal}]
    system = ("You are a sales-ops assistant. Use tools to gather facts before answering. "
              "If data is missing, say so. Finish with a 3-bullet briefing.")
    for step in range(max_steps):
        t0 = time.time()
        resp = client.messages.create(model=MODEL, max_tokens=4000, system=system,
                                      tools=TOOLS, messages=messages)
        log.info("step=%d stop=%s in=%d out=%d %.1fs", step, resp.stop_reason,
                 resp.usage.input_tokens, resp.usage.output_tokens, time.time() - t0)
        messages.append({"role": "assistant", "content": resp.content})
        if resp.stop_reason == "end_turn":
            return "".join(b.text for b in resp.content if b.type == "text")
        if resp.stop_reason == "max_tokens":
            messages.append({"role": "user", "content": "Continue, and be concise."})
            continue
        if resp.stop_reason != "tool_use":
            return f"Stopped: {resp.stop_reason}"
        results = []
        for block in resp.content:
            if block.type == "tool_use":
                output, is_error = run_tool(block.name, block.input)
                log.info("tool=%s args=%s error=%s", block.name, block.input, is_error)
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": output, "is_error": is_error})
        messages.append({"role": "user", "content": results})  # all results in ONE message
    return "Stopped: step budget exhausted"

if __name__ == "__main__":
    print(run_agent("Prepare me for a call with Acme Ltd tomorrow."))

Line-by-line: the parts that matter

  • Append the whole resp.content. It may include thinking blocks and multiple tool calls; dropping them breaks the conversation.
  • All tool results in one user message. Models can request several tools in parallel; splitting results across messages degrades that behavior.
  • Errors go back as is_error: true. The model can often recover (fix an argument, try another tool). Crashing the loop wastes the run.
  • Output caps. A tool that returns 200 KB of JSON will blow your context and your bill.
  • A step budget. Always. Pair it with a token or cost budget in production.
  • Structured logs per step. Stop reason, tokens, latency and tool names are the minimum trace you need (module 5 upgrades this to proper tracing).

Beyond the manual loop

The Anthropic SDKs include a beta tool runner (client.beta.messages.tool_runner in Python with the @beta_tool decorator) that runs this loop for you while still letting you inspect each turn. The Claude Agent SDK goes further, packaging Claude Code's harness (file, shell, web tools, hooks, sub-agents) as a library. Other ecosystems have equivalents: OpenAI's Agents SDK Runner, LangGraph's prebuilt agents, Google's ADK. Use them once you understand what they do for you.

Worked example trace

Goal: "Prepare me for a call with Acme Ltd tomorrow."

  1. Step 0: model calls search_crm(company="Acme Ltd").
  2. Step 1: sees Sara Khan, calls get_deals(contact_id="c_101").
  3. Step 2: end_turn with three bullets: who Sara is, the open £12,000 proposal, and a suggested agenda around the pricing call.

Three model calls, two tool calls. If your logs show twelve calls for this goal, something is wrong: vague tool descriptions, missing data, or instructions that never say when to stop.

Pitfalls

  • Parsing tool arguments with string matching instead of using the parsed input object.
  • Retrying non-retryable errors (400s) forever; let the SDK retry 429/5xx and fix the request for 4xx.
  • Letting tools raise exceptions that kill the process.
  • Forgetting that each loop iteration resends the entire history (costs grow roughly quadratically with steps unless you cache and trim).

Measuring success

For a fixed set of 20 goals, record: completion rate, average steps, average tool errors, tokens and cost per goal. These baseline numbers are what every later improvement is measured against.

Key takeaways

  • The loop: call the model, run requested tools, append results, repeat until end_turn or a budget stops it.
  • Append the assistant's full content and return all tool results in a single user message.
  • Return tool failures as is_error results so the model can recover instead of crashing.
  • Cap steps, tool output size and cost from the first version.
  • Log stop reason, tokens, latency and tool names for every step.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. The model returns stop_reason 'tool_use' with two tool_use blocks. How should you return the results?
  2. A tool throws an exception because an ID was malformed. What is the best behavior?
  3. Why does cost grow faster than linearly as an agent takes more steps?

Put it into practice

Run the loop in this lesson against a stub of one of your own tools. Record steps, tokens and latency for five goals and note one improvement.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.