---
title: "The explore-plan-implement-verify loop"
description: "Why agents need a loop, not a prompt The most common failure with coding agents is not bad code. It is the right code for the wrong plan , or a sprawling…"
url: https://optimizeall.com/learn/agentic-coding-with-ai/plan-implement-verify-loop
updated: 2026-10-05
---

AI-Assisted Software Development: Coding Agents in Practice · The plan-implement-verify workflow · lesson 7 of 17 · 14 min

# The explore-plan-implement-verify loop

## Why agents need a loop, not a prompt

The most common failure with coding agents is not bad code. It is **the right code for the wrong plan**, or a sprawling diff that nobody can review. The antidote is a deliberate loop with checkpoints where you, the engineer, apply judgment:

1. **Explore.** The agent reads relevant code and reports what it found. No edits.
2. **Plan.** The agent proposes a step-by-step plan, files to change, risks and a test strategy. You review and edit the plan.
3. **Implement.** The agent executes one step at a time in small, reviewable increments.
4. **Verify.** Tests, type checks, linters, and a manual check of behavior. The agent must show evidence, not claims.
5. **Review and commit.** You read the diff, then commit with a meaningful message. Repeat for the next step.

Anthropic's Claude Code best-practice guidance describes essentially this sequence (explore, plan, code, commit), and tools increasingly bake it in: Claude Code has a read-only **plan mode**; Cursor, Copilot and Codex support asking for a plan before edits. The mechanism differs, the discipline is the same.

## Checkpoint 1: exploration without edits

Ask explicitly for read-only exploration:

```text
Explore only; do not edit files yet. Read the code involved in CSV export and summarize:
(1) the request path from route to database, (2) where memory is allocated per row,
(3) existing tests that cover export, (4) anything surprising. Cite file paths and line ranges.
```

The citations let you spot-check claims in seconds. If the summary is wrong, the plan will be wrong; correct it now.

## Checkpoint 2: a plan worth reviewing

A good plan is short, ordered, and testable:

```markdown
Plan: stream CSV export
1. Add failing test: export of 10k fixture rows keeps heap growth under threshold (test/load/export.test.ts).
2. Add streamRows() in src/infra/db/stream.ts wrapper (reuse existing cursor helper).
3. Change exportCsv handler to pipe rows through csv-stringify stream; keep headers identical.
4. Byte-compare output with the existing small fixture.
5. Run npm test && npm run check; report results.
Risks: backpressure on slow clients; transaction held open during stream.
Out of scope: Excel export.
```

Your review questions: Is step order safe? Is the risk list honest? Does the test prove the goal? Edit the plan in the chat or in a `PLAN.md` file before any code is written. For large work, keep the plan in the repository so a fresh session (or a teammate) can resume.

## Checkpoint 3: small steps, frequent commits

Ask the agent to implement **one plan step at a time** and stop. Small diffs are easier to review and easier to revert. Commit after each verified step, so git becomes your undo button. Many teams use a dedicated branch per task, and for parallel agents, **git worktrees** (separate working directories on separate branches) so agents do not trample each other's files.

```bash
git worktree add ../invoicing-stream -b feat/stream-export
cd ../invoicing-stream && claude   # or codex, or open in your editor
```

## Checkpoint 4: verification with evidence

Require proof, not assertions:

```text
Run npm test and npm run check. Paste the final summary lines. If anything fails,
show the failure and your hypothesis before changing code.
```

Also verify what tests do not cover: run the app, hit the endpoint, look at the UI. For front-end work, some agents can drive a browser via MCP and capture screenshots; still look yourself before merging.

## Worked example: the loop in practice

A two-person studio in Manchester adds a "resend receipt" button to a Shopify-connected dashboard. Exploration reveals receipts are generated in two places (a legacy path and a new one). Without the exploration checkpoint, the agent would have changed only the new path. The plan adds a failing test for both paths, extracts a shared function, then wires the button. Four small commits; total review time under twenty minutes.

## When to break the loop

For trivial changes (rename a variable, fix a typo), the full loop is overkill; use pair mode and move on. For high-risk changes (auth, payments, data deletion), tighten the loop: plan reviewed by a second engineer, and no auto-accept of edits.

## Pitfalls

- **Skipping the plan** because the agent "seems to know". Plans are cheap; rework is not.
- **Accepting "all tests pass" without seeing output.** Agents occasionally misread results or run the wrong command.
- **Letting the agent edit tests to make them pass.** Watch for changed assertions in the diff.
- **Mega-commits.** If the diff is too big to review, it is too big to merge.

## How to measure success

Track average PR size for agent-assisted work and the share of agent PRs merged without major rework. Both should improve as the loop becomes habit.

## Video lecture: The explore-plan-implement-verify loop

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Plan → implement → verify
2. Analogy: renovating a kitchen
3. 1. Explore (read-only)
4. 2. Plan (you review it)
5. Persist big plans
6. 3. Implement in small steps
7. 4. Verify with evidence
8. 5. Review and commit
9. Scale the loop to risk
10. Example: 'last login' timestamp
11. Scenario: before vs after the loop (illustrative)
12. Deeper: two receipt paths
13. Watch me do it: streaming export
14. Recap

## Lecture transcript

### Plan → implement → verify

The most expensive mistake with a coding agent is not bad code. It is good code built on the wrong plan, delivered as a diff too big to review. In this lecture you will learn a five-step loop that professional teams use to prevent exactly that: explore, plan, implement, verify, and review. You will know what to check at each checkpoint and how to keep every step small enough to trust.

### Analogy: renovating a kitchen

Why a loop and not just a good prompt? Here is an analogy. Think of renovating a kitchen. A good contractor does not walk in and start knocking down walls. They survey the room, show you a drawing, agree it with you, then work one wall at a time, checking level and plumbing as they go. If you skip the drawing, you discover the sink is on the wrong side after the tiles are laid. The explore, plan, implement, verify loop is that renovation discipline, applied to code, with you as the client who signs off the drawing.

### 1. Explore (read-only)

Step one is explore, with no edits allowed. Ask the agent to read the relevant code and report back: the path a request takes, where the risky logic sits, which tests exist, and anything surprising. Insist on file paths and line ranges so you can spot-check its claims in seconds. If its understanding is wrong here, everything that follows will be wrong, and this is the cheapest possible moment to correct it.

### 2. Plan (you review it)

Step two is plan. A good plan is short, ordered and testable. It names the files to change, starts with a failing test that proves the goal, lists honest risks and states what is out of scope. Now you do the most valuable five minutes of the whole task. You review the plan. Is the order safe? Is the test meaningful? Are the risks real? Edit it before any code exists. Many tools help here: Claude Code has a read-only plan mode, and other agents will happily plan first if you ask.

### Persist big plans

For bigger work, keep the plan in a file in the repository, such as PLAN dot M D. That way a fresh session, or a teammate, can pick up exactly where you left off, without relying on a long chat history that may have been compacted. It also gives reviewers a clear statement of intent to compare the final diff against.

### 3. Implement in small steps

Step three is implement, one plan step at a time. Ask the agent to do step one and stop. Small diffs are easy to review and easy to revert. Commit after each verified step, and git becomes your undo button. If you run agents in parallel, give each its own branch and its own working directory using git worktrees, so they never trample each other's files.

### 4. Verify with evidence

Step four is verify, and the rule is evidence, not claims. Ask the agent to run the tests and type checks and paste the final summary lines. If something fails, it should show the failure and its hypothesis before touching code. Then verify what tests cannot: run the app, hit the endpoint, look at the screen. And scan the diff for one specific red flag: changed assertions in test files. An agent that edits the test to make it pass has not fixed anything.

### 5. Review and commit

Step five is review and commit. You read the diff as if a colleague wrote it, because in a sense one did. Then the loop repeats for the next step. Here is how it paid off for a two-person studio in Manchester adding a resend receipt button. Exploration revealed receipts were generated in two places, a legacy path and a new one. Without that checkpoint, the agent would have changed only one. The plan added tests for both, extracted a shared function, then wired the button. Four small commits. Under twenty minutes of review.

### Scale the loop to risk

The loop scales with risk. For a typo or a rename, skip it and just pair. For authentication, payments or data deletion, tighten it: a second engineer reviews the plan, and edits are never auto-accepted. The point is not ceremony. It is placing your judgment where it has the most leverage, before code exists and before code merges.

### Example: 'last login' timestamp

A simple example of the loop in action. Task: add a last login timestamp to the user profile page. Explore: the agent finds the user model, the login handler and the profile template, and notes there is already a migration pattern. Plan: one, add a failing test that logging in sets the timestamp; two, add the column with a migration; three, update the login handler; four, show it on the profile; five, run everything. You read the plan and add one line: store it in UTC and format it in the user's timezone. That one edit, made before any code, saves a round of review later.

### Scenario: before vs after the loop (illustrative)

Now a realistic scenario with illustrative numbers. A team in Islamabad compares two weeks of agent work before and after adopting the loop. Before: average agent pull request touched around six hundred lines, and about four in ten needed a second full review round. After: average size dropped to around one hundred and eighty lines, split into small commits, and only about one in ten needed a second round. Total time per feature barely changed, but reviewers stopped dreading agent pull requests. The common mistake to avoid: letting the agent run all plan steps in one go because it seems faster. It usually is not, once review is counted.

### Deeper: two receipt paths

One level deeper on the Manchester receipt example. The exploration found receipts generated in a legacy path used by the mobile app and a new path used by the web dashboard. The plan therefore started with two failing tests, one per path, then extracted a shared receipt builder, then wired the button to it. Without that exploration, the mobile app would have kept sending old-format receipts.

### Watch me do it: streaming export

Watch me do it with the streaming export task. Step one, explore: I paste the exploration prompt and tell the agent not to edit. It reports the route, the handler that builds an array of every row, the existing cursor helper, and the two tests covering export, all with file and line ranges. I open one citation to spot-check it. Correct. Step two, plan: it proposes five steps, starting with a failing load test. I read the risk list and add one it missed: a transaction held open while a slow client downloads. I move that into the plan as step three, use a read-only snapshot, and save the plan as PLAN dot M D. Step three, implement: I ask for step one only. It writes the load test, runs it, and shows it failing because memory grows with row count. I commit. Step two: the stream wrapper. Tests pass, I skim a thirty-line diff, commit. Step three: the handler change. Step four: the byte comparison with the fixture, which passes. Step five: the full suite plus type checks, with pasted summary lines. Finally, I scan the diff for edited test assertions. None. Five commits, each small enough to understand in a minute, and a plan file that explains why every change exists.

### Recap

Recap. Explore without edits. Plan, and review the plan. Implement one step at a time with commits. Verify with evidence, and watch for edited tests. Then review the diff. Your next step: take one real ticket this week and run the full loop, using the exploration and verification prompts in the lesson, and note how big your final pull request was compared with your usual.

## Key takeaways

- Explore read-only first and insist on file and line citations.
- Review and edit the plan before code exists; persist large plans in the repo.
- Implement one step at a time with commits; use git worktrees for parallel agents.
- Verify with evidence, not claims, and watch for edited test assertions.

## Try it

Run the full five-step loop on one real ticket, saving the plan in PLAN.md, and compare the final PR size with your usual.

- [Previous: Writing tasks agents can finish](https://optimizeall.com/learn/agentic-coding-with-ai/writing-tasks-agents-can-finish)
- [Next: Tests as guardrails: TDD with agents](https://optimizeall.com/learn/agentic-coding-with-ai/tests-as-guardrails)
- [All lessons of AI-Assisted Software Development: Coding Agents in Practice](https://optimizeall.com/learn/agentic-coding-with-ai)
