---
title: "How coding agents work: loops, context and permissions"
description: "The agent loop, demystified Every coding agent, whatever the brand, runs a variation of the same loop: 1. Gather context. Read the task, instruction…"
url: https://optimizeall.com/learn/agentic-coding-with-ai/how-coding-agents-work
updated: 2026-10-05
---

AI-Assisted Software Development: Coding Agents in Practice · The coding-agent landscape and how agents work · lesson 2 of 17 · 13 min

# How coding agents work: loops, context and permissions

## The agent loop, demystified

Every coding agent, whatever the brand, runs a variation of the same loop:

1. **Gather context.** Read the task, instruction files (CLAUDE.md, AGENTS.md), and whatever files the agent decides to open or search.
2. **Plan.** Decide on next steps; stronger agents write an explicit plan or to-do list.
3. **Act.** Call a tool: read a file, run `grep`, edit code, execute a shell command, query an MCP server.
4. **Observe.** Read the tool output: compiler errors, test results, diff.
5. **Repeat** until the model decides the task is complete, it hits a limit, or it asks you for input.

The model itself is stateless. Everything it "knows" about your project in a given moment is what sits in its **context window**: system prompt, instruction files, the conversation, and every tool result so far. That single fact explains most agent behavior, good and bad.

## Why context is the scarce resource

Context windows are large, but not free:

- **Cost and latency** grow with every token re-sent on each turn (prompt caching reduces but does not remove this).
- **Attention dilutes.** A window stuffed with irrelevant logs makes the model more likely to miss the one instruction that mattered.
- **Compaction loses detail.** When a long session fills the window, tools summarize earlier turns. Nuance from turn 5 may be gone by turn 80.

Practical consequence: short, focused sessions with a clear goal beat marathon sessions. Start a fresh session per task, and use subagents (separate context windows) for noisy exploration.

## Tools and permissions

An agent is only as dangerous as its tools. Typical built-in tools:

| Tool class | Examples | Risk |
|---|---|---|
| Read | open file, list directory, search | Low (but can read secrets) |
| Write | edit file, create file | Medium, reversible with git |
| Execute | shell commands, test runners, package installs | High: arbitrary code, network |
| External | MCP servers, web fetch, GitHub API | High: data leaves or enters your trust boundary |

Vendors expose permission controls in different shapes, but the concepts are shared:

- **Approval modes.** Ask before every edit or command; auto-accept edits but ask for commands; or fully autonomous. Codex exposes explicit sandbox and approval settings; Claude Code has permission modes (including a read-only plan mode) and allow/deny rules; Copilot's agent mode asks before running terminal commands unless you allow them.
- **Allow and deny lists.** Pre-approve safe commands (`npm test`, `pytest -q`) and block dangerous ones (`rm -rf`, `git push --force`, `curl` to arbitrary hosts).
- **Sandboxes.** Run the agent in a container, VM, cloud environment or OS-level sandbox with limited filesystem and network access.
- **Hooks.** Deterministic scripts that run before or after tool calls. Instruction files are *advice* to the model; hooks are *enforcement*. For example, Claude Code documents that CLAUDE.md is context, not enforced configuration, and recommends a PreToolUse hook when an action must be blocked regardless of what the model decides.

## Worked example: reading an agent transcript

Here is a condensed transcript of an agent fixing a failing test (tool calls in brackets):

```text
User: The test test_discount_rounding fails. Fix it without changing the public API.
Agent: I'll run the failing test first.       [run: pytest tests/test_pricing.py -k rounding]
  -> AssertionError: expected 19.99, got 19.989999
Agent: Looks like float arithmetic. Let me read pricing.py.   [read: src/pricing.py]
Agent: apply_discount uses float multiplication. I'll switch to Decimal with ROUND_HALF_UP.
                                                [edit: src/pricing.py]
                                                [run: pytest -q]
  -> 142 passed
Agent: Fixed by using Decimal in apply_discount; all 142 tests pass.
```

Notice what made this safe: the agent reproduced the failure first, made a small change, and verified with the full suite. Notice also what you still need to check: did switching to `Decimal` change the return type seen by callers? Tests passing is necessary, not sufficient.

## Hands-on: set a sane permission baseline

A minimal Claude Code project settings file that allows common read-only and test commands and denies risky ones (check the current settings reference for exact syntax):

```json
{
  "permissions": {
    "allow": ["Bash(npm test:*)", "Bash(npm run lint)", "Bash(git diff:*)", "Bash(git status)"],
    "deny": ["Bash(git push:*)", "Bash(rm -rf:*)", "Read(./.env)", "Read(./.env.*)", "Read(./secrets/**)"]
  }
}
```

The equivalent idea in Codex is to start in a restrictive sandbox and approval mode and widen only when needed; in Copilot agent mode, review the terminal-command approval settings in your editor. Whatever the tool, write down your team's baseline and commit it to the repository so everyone starts from the same safe defaults.

## Pitfalls

- **Assuming the agent "remembers".** It remembers only what is in context or in a memory file.
- **Letting sessions sprawl.** After a long session, start fresh with a short summary of state.
- **Relying on instructions for safety.** "Never push to main" in a Markdown file is a request. Branch protection and hooks are controls.

## How to measure success

Track two signals for a sprint: how often an agent run needed a restart because it lost the thread (aim to reduce by shortening sessions), and how many permission prompts were for commands you always approve (move those to the allowlist).

## Video lecture: How coding agents work: loops, context and permissions

Lecture coming soon · 13 chapters · about 9 minutes. Read the full transcript below.

1. How coding agents actually work
2. Analogy: the contractor's whiteboard
3. The agent loop
4. Context is everything
5. Tool risk ladder
6. Four controls
7. Reading a transcript
8. Where the agent runs
9. Example: 'always use pnpm'
10. Scenario: the accidental leak
11. Deeper: what the tests didn't cover
12. Watch me do it: the permission baseline
13. Recap

## Lecture transcript

### How coding agents actually work

If you have ever watched a coding agent do something brilliant and then, ten minutes later, forget a rule you told it at the start, this lecture explains why. By the end you will understand the agent loop, why context is the scarce resource, and how permissions, sandboxes and hooks keep agents safe.

### Analogy: the contractor's whiteboard

Before we go deeper, here is an analogy that makes the rest of this lecture click. Think of the agent as a very fast contractor working at a desk with one small whiteboard. Everything it knows about your project has to fit on that whiteboard: your instructions, the files it has read, the output of every command. When the whiteboard fills up, someone wipes the oldest notes and writes a short summary in the corner. The contractor is brilliant, but it can only work with what is on the whiteboard right now. Keep that picture in mind: the whiteboard is the context window, and the wiping is compaction.

### The agent loop

Every coding agent runs the same basic loop. It gathers context: your request, instruction files and whatever files it chooses to read. It plans. It acts by calling a tool, like reading a file, running a search, editing code or executing a command. It observes the result, like a compiler error or a test report. Then it repeats until it believes the job is done, hits a limit, or needs your input. The brand on the box changes the quality of each step, but not the shape of the loop.

### Context is everything

Here is the most important fact about agents. The model is stateless. In any given moment, everything it knows about your project is what sits in its context window. The system prompt, your instruction file, the conversation, and every tool result so far. When a long session fills that window, the tool compacts older turns into a summary, and detail gets lost. That is why the rule you gave at turn five can vanish by turn eighty. The fix is simple: short, focused sessions with one clear goal, and fresh sessions for new tasks.

### Tool risk ladder

Next, tools. Reading files is low risk, although an agent can read secrets if you let it. Writing files is medium risk, and git makes it reversible. Executing commands is high risk. A shell command can do anything, including installing packages or calling the network. And external tools, like MCP servers or web fetch, move data across your trust boundary in both directions. So the question is never just which agent. It is which tools, with which permissions.

### Four controls

Vendors expose permissions differently, but the concepts are the same. Approval modes decide whether the agent asks before every edit or command. Allow and deny lists pre-approve safe commands like running tests, and block dangerous ones like force pushes. Sandboxes restrict what files and network the agent can touch. And hooks are deterministic scripts that run before or after a tool call. Here is the crucial distinction. Instructions in a Markdown file are advice to the model. Hooks, sandboxes and branch protection are enforcement. Anthropic's own docs make exactly this point about its memory file.

### Reading a transcript

Let's read a short transcript. The agent is asked to fix a failing rounding test without changing the public API. First it runs the failing test, and sees nineteen point nine eight nine nine instead of nineteen point nine nine. It reads the pricing module, spots floating point multiplication, switches to a decimal type, and reruns the whole suite. All green. That is good agent behavior: reproduce, small change, full verification. But your job is not over. Did the return type change for callers? Passing tests are necessary. They are not sufficient.

### Where the agent runs

One more design choice you control is where the agent runs. A local agent in your terminal inherits your permissions: your SSH keys, your cloud credentials, your browser cookies if a tool can reach them. A cloud agent runs in an isolated environment the vendor provisions, usually with a fresh checkout, limited network access and a scoped token. Neither is automatically safe. Local agents need allow lists and sandboxing; cloud agents need tight repository scopes and firewall rules. Also use subagents, which are separate context windows, for noisy jobs like searching a huge codebase, so the main session stays clean and focused on the task.

### Example: 'always use pnpm'

A simple example to lock it in. Say you tell the agent at the start, always use pnpm, never npm. An hour later, it runs npm install. Why? Look at the whiteboard. Your instruction was one short line in the chat, buried under forty tool outputs, and compaction summarized it away. The fix takes ten seconds: put use pnpm in your instruction file, which is loaded fresh into every session, and start a new session for the next task. And if running npm install would actually cause damage, add a deny rule so the command is blocked no matter what the agent decides. Advice goes in the file. Enforcement goes in permissions.

### Scenario: the accidental leak

Now a realistic scenario. A fintech team in Karachi lets an agent run with auto-approved commands on a developer laptop. The laptop also holds cloud credentials in its environment. A test script the agent runs happens to print environment variables when it fails, and the agent helpfully pastes that output into a pull request description. Nobody intended a leak, yet a secret ends up in a pull request. What would have stopped it? Denying the agent read access to environment files, running agents in a sandbox without production credentials, and secret scanning on pull requests. Three independent layers, any one of which would have caught it. That is the mindset: assume one layer fails, and make sure another catches it.

### Deeper: what the tests didn't cover

One level deeper on the transcript example. After the agent switched to a decimal type, the suite passed. But the function used to return a float, and two callers in the reporting module compare it with plain numbers. A quick search for the function name shows those callers, so you ask the agent to add one test per caller before merging.

### Watch me do it: the permission baseline

Watch me do it. Here is the permission baseline from the lesson, and I'll walk through it line by line so you know why each line is there. The allow list starts with Bash npm test colon star. That means any command starting with npm test runs without asking me, because tests are the agent's feedback loop, and I will approve them a hundred times a day anyway. Next, npm run lint, same reason. Then git diff and git status, read-only commands the agent uses to understand what it changed. Now the deny list. Bash git push colon star: the agent can commit locally, but pushing is a human decision, and branch protection backs this up on the server. Bash rm dash rf: no recursive deletes, ever. Then three Read rules: the env file, anything starting with dot env, and the secrets folder. Even reading those puts credentials into the context window, where they can leak. Now let me test it. I start a session and ask the agent to print the contents of the env file. The tool call is refused before it runs, and the agent tells me it is not permitted. Then I ask it to run the tests, and they run with no prompt at all. That is what a good baseline feels like: routine work flows, dangerous work stops. Commit this file, and every developer gets the same guardrails.

### Recap

To recap. Agents loop through context, plan, act and observe. Context is the scarce resource, so keep sessions short and focused. Tool risk rises from read to write to execute to external. And real safety comes from enforcement, not instructions. Your next step: copy the permission baseline from the lesson, adapt it to your stack, and commit it to your repository so the whole team starts from the same safe defaults.

## Key takeaways

- Agents loop: gather context, plan, act with tools, observe, repeat.
- The model is stateless; only what is in the context window counts, and compaction loses detail.
- Tool risk rises from read to write to execute to external systems.
- Instruction files are advice; hooks, sandboxes and branch protection are enforcement.

## Try it

Write a permission baseline for your main repository (allow, deny, protected paths) and commit it on a branch for review.

- [Previous: From autocomplete to agents: the 2026 landscape](https://optimizeall.com/learn/agentic-coding-with-ai/from-autocomplete-to-agents)
- [Next: Choosing and setting up tools with a fair trial](https://optimizeall.com/learn/agentic-coding-with-ai/choosing-and-setting-up-tools)
- [All lessons of AI-Assisted Software Development: Coding Agents in Practice](https://optimizeall.com/learn/agentic-coding-with-ai)
