AI-Assisted Software Development: Coding Agents in PracticeThe coding-agent landscape and how agents work · Lesson 2 of 17

How coding agents work: loops, context and permissions

Article · 13 min · 9 min lecture

Video lecture

How coding agents work: loops, context and permissions

13 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 13

How coding agents actually work

  • The agent loop
  • Context: the scarce resource
  • Tools, permissions, sandboxes, hooks

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The agent loop, demystified

Every coding agent, whatever the brand, runs a variation of the same loop:

  1. Gather context. Read the task, instruction files (CLAUDE.md, AGENTS.md), and whatever files the agent decides to open or search.
  2. Plan. Decide on next steps; stronger agents write an explicit plan or to-do list.
  3. Act. Call a tool: read a file, run grep, edit code, execute a shell command, query an MCP server.
  4. Observe. Read the tool output: compiler errors, test results, diff.
  5. Repeat until the model decides the task is complete, it hits a limit, or it asks you for input.

The model itself is stateless. Everything it "knows" about your project in a given moment is what sits in its context window: system prompt, instruction files, the conversation, and every tool result so far. That single fact explains most agent behavior, good and bad.

Why context is the scarce resource

Context windows are large, but not free:

  • Cost and latency grow with every token re-sent on each turn (prompt caching reduces but does not remove this).
  • Attention dilutes. A window stuffed with irrelevant logs makes the model more likely to miss the one instruction that mattered.
  • Compaction loses detail. When a long session fills the window, tools summarize earlier turns. Nuance from turn 5 may be gone by turn 80.

Practical consequence: short, focused sessions with a clear goal beat marathon sessions. Start a fresh session per task, and use subagents (separate context windows) for noisy exploration.

Tools and permissions

An agent is only as dangerous as its tools. Typical built-in tools:

Tool classExamplesRisk
Readopen file, list directory, searchLow (but can read secrets)
Writeedit file, create fileMedium, reversible with git
Executeshell commands, test runners, package installsHigh: arbitrary code, network
ExternalMCP servers, web fetch, GitHub APIHigh: data leaves or enters your trust boundary

Vendors expose permission controls in different shapes, but the concepts are shared:

  • Approval modes. Ask before every edit or command; auto-accept edits but ask for commands; or fully autonomous. Codex exposes explicit sandbox and approval settings; Claude Code has permission modes (including a read-only plan mode) and allow/deny rules; Copilot's agent mode asks before running terminal commands unless you allow them.
  • Allow and deny lists. Pre-approve safe commands (npm test, pytest -q) and block dangerous ones (rm -rf, git push --force, curl to arbitrary hosts).
  • Sandboxes. Run the agent in a container, VM, cloud environment or OS-level sandbox with limited filesystem and network access.
  • Hooks. Deterministic scripts that run before or after tool calls. Instruction files are advice to the model; hooks are enforcement. For example, Claude Code documents that CLAUDE.md is context, not enforced configuration, and recommends a PreToolUse hook when an action must be blocked regardless of what the model decides.

Worked example: reading an agent transcript

Here is a condensed transcript of an agent fixing a failing test (tool calls in brackets):

User: The test test_discount_rounding fails. Fix it without changing the public API.
Agent: I'll run the failing test first.       [run: pytest tests/test_pricing.py -k rounding]
  -> AssertionError: expected 19.99, got 19.989999
Agent: Looks like float arithmetic. Let me read pricing.py.   [read: src/pricing.py]
Agent: apply_discount uses float multiplication. I'll switch to Decimal with ROUND_HALF_UP.
                                                [edit: src/pricing.py]
                                                [run: pytest -q]
  -> 142 passed
Agent: Fixed by using Decimal in apply_discount; all 142 tests pass.

Notice what made this safe: the agent reproduced the failure first, made a small change, and verified with the full suite. Notice also what you still need to check: did switching to Decimal change the return type seen by callers? Tests passing is necessary, not sufficient.

Hands-on: set a sane permission baseline

A minimal Claude Code project settings file that allows common read-only and test commands and denies risky ones (check the current settings reference for exact syntax):

{
  "permissions": {
    "allow": ["Bash(npm test:*)", "Bash(npm run lint)", "Bash(git diff:*)", "Bash(git status)"],
    "deny": ["Bash(git push:*)", "Bash(rm -rf:*)", "Read(./.env)", "Read(./.env.*)", "Read(./secrets/**)"]
  }
}

The equivalent idea in Codex is to start in a restrictive sandbox and approval mode and widen only when needed; in Copilot agent mode, review the terminal-command approval settings in your editor. Whatever the tool, write down your team's baseline and commit it to the repository so everyone starts from the same safe defaults.

Pitfalls

  • Assuming the agent "remembers". It remembers only what is in context or in a memory file.
  • Letting sessions sprawl. After a long session, start fresh with a short summary of state.
  • Relying on instructions for safety. "Never push to main" in a Markdown file is a request. Branch protection and hooks are controls.

How to measure success

Track two signals for a sprint: how often an agent run needed a restart because it lost the thread (aim to reduce by shortening sessions), and how many permission prompts were for commands you always approve (move those to the allowlist).

Key takeaways

  • Agents loop: gather context, plan, act with tools, observe, repeat.
  • The model is stateless; only what is in the context window counts, and compaction loses detail.
  • Tool risk rises from read to write to execute to external systems.
  • Instruction files are advice; hooks, sandboxes and branch protection are enforcement.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent ignores a rule you gave at the start of a very long session. What is the most likely cause?
  2. You must guarantee an agent can never run `git push --force`. What is the right control?

Put it into practice

Write a permission baseline for your main repository (allow, deny, protected paths) and commit it on a branch for review.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.