---
title: "Build the agent loop from scratch in Python"
description: "Why build it by hand first Frameworks and SDK helpers are great, but if you have never written the loop yourself you will struggle to debug them. The raw…"
url: https://optimizeall.com/learn/ai-agents-engineering/agent-loop-from-scratch
updated: 2026-10-05
---

Building Production AI Agents · Agent foundations and architectures · lesson 3 of 18 · 18 min

# Build the agent loop from scratch in Python

## Why build it by hand first

Frameworks and SDK helpers are great, but if you have never written the loop yourself you will struggle to debug them. The raw loop is about 40 lines. Once you understand it, every framework is just conveniences around it.

## The anatomy of a tool call (Claude Messages API)

1. You send `messages`, a `system` prompt and a list of `tools`, each with a `name`, `description` and JSON Schema `input_schema`.
2. The model replies. If it wants a tool, `stop_reason` is `"tool_use"` and the `content` contains one or more `tool_use` blocks with an `id`, `name` and `input`.
3. You append the assistant's **full** `content` (not just its text) to `messages`, run each tool, then append **one** user message containing a `tool_result` block per call, each referencing the matching `tool_use_id`.
4. Call the API again. Repeat until `stop_reason` is `"end_turn"` (or something else you must handle).

Other stop reasons to handle: `"max_tokens"` (output was cut off; raise the limit or ask it to continue), `"pause_turn"` (a long server-side tool turn paused; send the conversation back to continue), and `"refusal"` on some newer models (the model declined; do not read tool calls from it). OpenAI's Responses API and Google's Gemini API follow the same idea with different field names, which the platform-integration course compares side by side.

## Hands-on: a minimal but production-minded loop

```python
import os, json, time, logging
import anthropic

logging.basicConfig(level=logging.INFO)
log = logging.getLogger("agent")
client = anthropic.Anthropic()  # ANTHROPIC_API_KEY from env; SDK retries 429/5xx by default
MODEL = os.environ.get("AGENT_MODEL", "claude-sonnet-5")  # verify current IDs in the docs

TOOLS = [
    {
        "name": "search_crm",
        "description": "Search CRM contacts by company name. Returns up to 5 matches with id, name, role, last_contacted.",
        "input_schema": {
            "type": "object",
            "properties": {"company": {"type": "string", "description": "Company name, e.g. 'Acme Ltd'"}},
            "required": ["company"],
            "additionalProperties": False,
        },
    },
    {
        "name": "get_deals",
        "description": "List open deals for a CRM contact id. Returns stage, value and next step.",
        "input_schema": {
            "type": "object",
            "properties": {"contact_id": {"type": "string"}},
            "required": ["contact_id"],
            "additionalProperties": False,
        },
    },
]

def search_crm(company: str) -> list[dict]:        # replace with your real CRM client
    return [{"id": "c_101", "name": "Sara Khan", "role": "Head of Growth", "last_contacted": "2026-08-30"}]

def get_deals(contact_id: str) -> list[dict]:
    return [{"stage": "proposal", "value_gbp": 12000, "next_step": "pricing call"}]

REGISTRY = {"search_crm": search_crm, "get_deals": get_deals}

def run_tool(name: str, args: dict) -> tuple[str, bool]:
    fn = REGISTRY.get(name)
    if fn is None:
        return f"Unknown tool {name}", True
    try:
        return json.dumps(fn(**args))[:8000], False   # cap tool output size
    except Exception as exc:                          # return errors to the model, don't crash
        return f"Tool error: {type(exc).__name__}: {exc}", True

def run_agent(goal: str, max_steps: int = 10) -> str:
    messages = [{"role": "user", "content": goal}]
    system = ("You are a sales-ops assistant. Use tools to gather facts before answering. "
              "If data is missing, say so. Finish with a 3-bullet briefing.")
    for step in range(max_steps):
        t0 = time.time()
        resp = client.messages.create(model=MODEL, max_tokens=4000, system=system,
                                      tools=TOOLS, messages=messages)
        log.info("step=%d stop=%s in=%d out=%d %.1fs", step, resp.stop_reason,
                 resp.usage.input_tokens, resp.usage.output_tokens, time.time() - t0)
        messages.append({"role": "assistant", "content": resp.content})
        if resp.stop_reason == "end_turn":
            return "".join(b.text for b in resp.content if b.type == "text")
        if resp.stop_reason == "max_tokens":
            messages.append({"role": "user", "content": "Continue, and be concise."})
            continue
        if resp.stop_reason != "tool_use":
            return f"Stopped: {resp.stop_reason}"
        results = []
        for block in resp.content:
            if block.type == "tool_use":
                output, is_error = run_tool(block.name, block.input)
                log.info("tool=%s args=%s error=%s", block.name, block.input, is_error)
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": output, "is_error": is_error})
        messages.append({"role": "user", "content": results})  # all results in ONE message
    return "Stopped: step budget exhausted"

if __name__ == "__main__":
    print(run_agent("Prepare me for a call with Acme Ltd tomorrow."))
```

## Line-by-line: the parts that matter

- **Append the whole `resp.content`.** It may include thinking blocks and multiple tool calls; dropping them breaks the conversation.
- **All tool results in one user message.** Models can request several tools in parallel; splitting results across messages degrades that behavior.
- **Errors go back as `is_error: true`.** The model can often recover (fix an argument, try another tool). Crashing the loop wastes the run.
- **Output caps.** A tool that returns 200 KB of JSON will blow your context and your bill.
- **A step budget.** Always. Pair it with a token or cost budget in production.
- **Structured logs per step.** Stop reason, tokens, latency and tool names are the minimum trace you need (module 5 upgrades this to proper tracing).

## Beyond the manual loop

The Anthropic SDKs include a beta **tool runner** (`client.beta.messages.tool_runner` in Python with the `@beta_tool` decorator) that runs this loop for you while still letting you inspect each turn. The **Claude Agent SDK** goes further, packaging Claude Code's harness (file, shell, web tools, hooks, sub-agents) as a library. Other ecosystems have equivalents: OpenAI's Agents SDK `Runner`, LangGraph's prebuilt agents, Google's ADK. Use them once you understand what they do for you.

## Worked example trace

Goal: "Prepare me for a call with Acme Ltd tomorrow."

1. Step 0: model calls `search_crm(company="Acme Ltd")`.
2. Step 1: sees Sara Khan, calls `get_deals(contact_id="c_101")`.
3. Step 2: `end_turn` with three bullets: who Sara is, the open £12,000 proposal, and a suggested agenda around the pricing call.

Three model calls, two tool calls. If your logs show twelve calls for this goal, something is wrong: vague tool descriptions, missing data, or instructions that never say when to stop.

## Pitfalls

- Parsing tool arguments with string matching instead of using the parsed `input` object.
- Retrying non-retryable errors (400s) forever; let the SDK retry 429/5xx and fix the request for 4xx.
- Letting tools raise exceptions that kill the process.
- Forgetting that each loop iteration resends the entire history (costs grow roughly quadratically with steps unless you cache and trim).

## Measuring success

For a fixed set of 20 goals, record: completion rate, average steps, average tool errors, tokens and cost per goal. These baseline numbers are what every later improvement is measured against.

## Video lecture: Build the agent loop from scratch in Python

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Build the agent loop from scratch
2. Why build it yourself?
3. Anatomy of a tool call
4. Two rules
5. Handle every stop reason
6. Simple trace: umbrella question
7. Production habits
8. Worked trace
9. Cost grows with steps
10. Errors: fragile vs robust
11. Deeper: the ambiguous-account bug
12. Watch me do it: run_agent()
13. Try this now
14. Recap

## Lecture transcript

### Build the agent loop from scratch

Every agent framework you'll ever use hides the same forty lines of code. If you've written those lines yourself, debugging any framework becomes easy. In this lesson you'll build the agent loop from scratch in Python, with the production habits that most tutorials skip.

### Why build it yourself?

Why build the loop yourself when frameworks exist? Because when something breaks at two in the morning, you need to know what's actually happening. Think of learning to drive a manual car before an automatic. You may never drive stick again, but you'll understand what the gearbox is doing. Every framework, whether it's an SDK tool runner, LangGraph or the Claude Agent SDK, runs this same cycle underneath. Once you've written it, framework docs stop being magic and start being a list of conveniences, and debugging becomes reading a transcript instead of guessing.

### Anatomy of a tool call

Start with the anatomy of a tool call. You send the model your messages, a system prompt, and a list of tools. Each tool has a name, a description and a JSON schema for its inputs. If the model wants a tool, the stop reason is tool use, and the content contains tool use blocks, each with an id, a name and input arguments. Your code runs the tool, then sends back tool result blocks that point to those ids. Then you call the model again.

### Two rules

Two rules trip people up. First, append the assistant's whole content to the history, not just the text. It can include thinking blocks and several tool calls. Second, when the model asks for multiple tools at once, return all the results together in one user message. Split them up and the model learns to stop calling tools in parallel, which slows everything down.

### Handle every stop reason

Now the stop reasons. End turn means the model is done. Tool use means run tools and loop. Max tokens means the output was cut off, so continue or raise the limit. Pause turn appears with long server side tools, and you just send the conversation back. And on some newer models, refusal means the model declined, so don't try to read tool calls from that response. Handle each one explicitly and your loop never gets stuck.

### Simple trace: umbrella question

Let's trace a simple example first, before the business one. You give the agent one tool, get weather, and ask: should I take an umbrella in Manchester tomorrow? Call one: the model replies with stop reason tool use and a tool block asking for get weather with city Manchester. Your code runs the function and gets back, rain likely, eighty percent. You append the assistant's content, then a user message with the tool result pointing to that block's id. Call two: the model replies with end turn: yes, take an umbrella. Two calls, one tool. That's the whole machine.

### Production habits

Let's talk production habits. Errors: if a tool fails, return the error message with is error set to true. The model often fixes its own mistake, like a malformed id. Size: cap tool outputs, because a two hundred kilobyte JSON blob will swamp the context and your bill. Budgets: always set a maximum number of steps, and in production add a token or cost budget. Logs: record the stop reason, tokens, latency and tool names for every step.

### Worked trace

Here's a worked trace. The goal: prepare me for a call with Acme tomorrow. Step zero, the model searches the CRM for Acme and finds Sara Khan, Head of Growth. Step one, it fetches her open deals: a twelve thousand pound proposal waiting on a pricing call. Step two, it ends the turn with three crisp bullets. Three model calls, two tool calls. If your logs show twelve calls for a goal like this, that's a signal: vague tool descriptions, missing data, or no clear finish line.

### Cost grows with steps

One cost insight before we finish. Every step resends the entire conversation, so input tokens pile up with each loop. That's why long agent runs get expensive fast, and why later lessons cover prompt caching and context trimming. Once you understand the manual loop, you can use helpers like the Anthropic tool runner, the Claude Agent SDK, or other frameworks, knowing exactly what they're doing for you.

### Errors: fragile vs robust

Let's zoom in on the error handling, because it's where amateur loops and professional loops really differ. Picture a tool that looks up an order, and the model passes an order number with a typo. A fragile loop throws an exception and the whole run dies, wasting every token spent so far. A robust loop catches the exception and returns a short, specific message: order not found, ids look like O R D dash five digits. Nine times out of ten, the model reads that, fixes the argument and carries on. Also separate two kinds of failure. Rate limits and server errors are retryable, and the official SDKs already retry them with backoff. Bad requests, like an invalid schema, are not retryable. Fix the request instead of hammering the API.

### Deeper: the ambiguous-account bug

Let's deepen the CRM briefing example. Imagine the same agent at a Karachi software reseller with two thousand CRM contacts. On a normal day, briefing requests take three model calls and two tool calls, as we saw. But when a company name matches several records, like two different Acme entities, the first version of the agent picked the first match and briefed the salesperson on the wrong account. The fix was not a smarter model. The team changed search CRM to return a city and industry with each match, and added one sentence to the system prompt: if more than one company matches, ask which one. Wrong account briefings dropped to almost zero, at the cost of an occasional clarifying question, which salespeople actually liked.

### Watch me do it: run_agent()

Watch me do it. Let's run the loop from the lesson and follow it line by line. First, the tools list: search CRM and get deals, each with a description and a strict input schema. Next, run tool looks up the function in the registry, calls it, caps the output at eight thousand characters, and turns any exception into an error string with is error true. Now run agent. It starts the messages with the goal, then loops up to ten steps. Each step calls messages create with the system prompt, tools and messages, and logs the step number, stop reason, tokens and time. It appends the whole assistant content. If the stop reason is end turn, it returns the text. If it's max tokens, it asks the model to continue concisely. Otherwise, for every tool use block, it runs the tool and collects a tool result with the matching id, then appends all results in one user message. When I run it, the log shows three steps: tool use, tool use, end turn.

### Try this now

Here's your try this now. Copy the loop from the lesson text and replace the two CRM functions with a stub of one tool you actually use, even if it just returns fake data. Run five different goals. For each, write down the number of steps, total input and output tokens, and seconds taken. Then deliberately break the tool: make it throw an error and watch whether the model recovers when it sees the error message. That ten minute experiment teaches you more about agent behavior than any diagram.

### Recap

To recap: the loop is call, run tools, append results, repeat. Append full content, return results together, send errors back as data, cap everything, and log every step. Your next step: run the lesson's code against a stub of one of your own tools, try five goals, and record the steps, tokens and latency. That's your baseline for every improvement to come.

## Key takeaways

- The loop: call the model, run requested tools, append results, repeat until end_turn or a budget stops it.
- Append the assistant's full content and return all tool results in a single user message.
- Return tool failures as is_error results so the model can recover instead of crashing.
- Cap steps, tool output size and cost from the first version.
- Log stop reason, tokens, latency and tool names for every step.

## Try it

Run the loop in this lesson against a stub of one of your own tools. Record steps, tokens and latency for five goals and note one improvement.

- [Previous: Workflow and agent architectures: the pattern catalog](https://optimizeall.com/learn/ai-agents-engineering/workflow-and-agent-patterns)
- [Next: Designing tools agents use correctly](https://optimizeall.com/learn/ai-agents-engineering/designing-agent-tools)
- [All lessons of Building Production AI Agents](https://optimizeall.com/learn/ai-agents-engineering)
