---
title: "Debugging with agents: a scientific protocol"
description: "Debugging is where agents shine, if you run it like science Debugging suits agents well: they can read stack traces, grep through code, add logging, and…"
url: https://optimizeall.com/learn/agentic-coding-with-ai/debugging-with-agents
updated: 2026-10-05
---

AI-Assisted Software Development: Coding Agents in Practice · The plan-implement-verify workflow · lesson 9 of 17 · 13 min

# Debugging with agents: a scientific protocol

## Debugging is where agents shine, if you run it like science

Debugging suits agents well: they can read stack traces, grep through code, add logging, and rerun a reproduction dozens of times without getting tired. They are also prone to a classic failure: **changing code until the symptom disappears** without understanding the cause. Your job is to impose the scientific method.

## The debugging protocol

1. **Reproduce.** Get a failing test or a script that reliably shows the bug. No reproduction, no fix.
2. **Gather evidence.** Logs, traces, inputs, recent changes (`git log -p`, `git bisect`).
3. **Hypothesize.** List candidate causes ranked by likelihood, each with a way to confirm or refute it.
4. **Test hypotheses.** Add targeted logging or assertions; run the reproduction; eliminate causes.
5. **Fix the cause**, not the symptom, with the smallest change.
6. **Prove it.** The reproduction now passes; the full suite passes; add a regression test.

Encode this as a reusable prompt or skill:

```text
Debug protocol. Do not change application code until step 5.
1) Reproduce the bug as a failing automated test and show me it fails.
2) Collect evidence: relevant logs, inputs, and recent commits touching these files.
3) List 3 hypotheses ranked by likelihood, each with a cheap experiment to test it.
4) Run the experiments (temporary logging is fine) and report which hypothesis survives.
5) Propose the minimal fix for the root cause and wait for my approval.
6) After approval: implement, remove temporary logging, run the full suite, keep the test.
```

## Using git to narrow the search

When "it worked last week", `git bisect` finds the commit that introduced a bug by binary search. Agents can run it automatically when you give them a test command:

```bash
git bisect start
git bisect bad HEAD
git bisect good v2.14.0
git bisect run npx vitest run test/export.test.ts
git bisect reset
```

Ask the agent to read the offending commit and explain *why* it broke the behavior before fixing.

## Debugging production issues safely

For production incidents, the agent should work from **copies of evidence**, not live systems:

- Export sanitized logs or traces into the repo's scratch folder; strip customer PII first.
- Give read-only access to observability tools through MCP only if the credentials are scoped and audited.
- Never let an agent run commands against production databases during an incident. Use a replica or snapshot.

## Worked example: the intermittent failure

A logistics startup in Islamabad saw a test that failed roughly one run in twenty in CI. A developer asked an agent to "fix the flaky test"; its first attempt added a retry, which hid the problem. Running the protocol instead:

- Reproduction: the agent ran the test 200 times in a loop and captured failing seeds.
- Evidence: failures occurred only when two jobs shared a timestamp.
- Hypotheses: (a) clock resolution, (b) unstable sort on equal timestamps, (c) database isolation.
- Experiment: forcing equal timestamps failed every time; the sort was not stable on ties.
- Fix: add a secondary sort key (job ID) in the query. Regression test pinned equal timestamps.

The retry would have shipped the real bug to customers, whose shipments occasionally appeared out of order.

## Reading unfamiliar errors and stacks

Agents are excellent at explaining unfamiliar errors, framework internals and cryptic stack traces. Ask for the explanation *and* the evidence: "Which line in our code is the last frame we control? What input would cause this path?" For library bugs, ask the agent to check the installed version's changelog or source in `node_modules` or site-packages rather than relying on memory, which may describe a different version.

## Hands-on: a loop runner for flaky tests

```bash
#!/usr/bin/env bash
# run-many.sh <n> <test command...>: run a test repeatedly and count failures
n=$1; shift; fails=0
for i in $(seq 1 "$n"); do
  if ! "$@" > /tmp/run-$i.log 2>&1; then fails=$((fails+1)); echo "fail on run $i"; fi
done
echo "$fails failures out of $n runs"
```

Give this to the agent as a tool for reproducing intermittent bugs.

## Pitfalls

- **Retry as a fix.** Retries hide real bugs unless the failure is truly external and transient.
- **Shotgun edits.** Several simultaneous changes make it impossible to know what fixed it.
- **Leftover debug logging** containing sensitive data. Make removal part of the protocol.
- **Trusting remembered APIs.** Versions differ; read the installed code.

## How to measure success

Track mean time to root cause, the share of bug fixes that include a regression test, and bugs that reopen within 30 days.

## Video lecture: Debugging with agents: a scientific protocol

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Debugging with agents
2. Analogy: the good doctor
3. Steps 1–3
4. Steps 4–6
5. git bisect
6. Case: the flaky test
7. Production debugging
8. Reading unfamiliar errors
9. Pitfalls
10. Example: blank search page
11. Scenario: 'random' checkout errors (illustrative)
12. Deeper: making 'random' deterministic
13. Watch me do it: the flaky test
14. Recap

## Lecture transcript

### Debugging with agents

Agents are tireless debuggers. They will read stack traces, add logging and rerun a reproduction two hundred times without complaint. They also share a very human bad habit: changing code until the symptom goes away, without understanding why. In this lecture you will learn a six-step debugging protocol that forces the scientific method, how to use git to narrow the search, and how to debug production issues safely.

### Analogy: the good doctor

Why does debugging need a protocol? Think about a good doctor. When you arrive with a headache, they do not immediately prescribe the strongest painkiller. They ask questions, take your temperature, rule things out, and only then treat the cause. An agent without a protocol behaves like a doctor who prescribes painkillers for everything: the symptom goes away and the illness stays. The protocol turns your agent into the good doctor: reproduce, examine, hypothesize, test, then treat the cause.

### Steps 1–3

Step one: reproduce. Get a failing test or a script that reliably shows the bug. No reproduction, no fix. Step two: gather evidence, like logs, inputs and recent commits touching the affected files. Step three: hypothesize. The agent lists several candidate causes ranked by likelihood, each with a cheap experiment to confirm or refute it. This is where you add the most value, by spotting the hypothesis the agent missed.

### Steps 4–6

Step four: test the hypotheses. Add temporary logging or assertions, run the reproduction, and eliminate causes one by one. Step five: fix the root cause with the smallest possible change, and here the agent should pause for your approval. Step six: prove it. The reproduction passes, the full suite passes, temporary logging is removed, and a regression test stays forever. The lesson has this whole protocol as a copy-paste prompt.

### git bisect

When something worked last week and not today, use git bisect. It performs a binary search through your history to find the commit that introduced the bug. Give it a test command and it runs automatically, which is perfect for an agent. Then ask the agent to read the offending commit and explain why it broke the behavior before it fixes anything. Understanding first, change second.

### Case: the flaky test

A story from a logistics startup in Islamabad. One test failed about once every twenty runs in CI. The first agent attempt added a retry, and the test went green. That would have shipped a real bug. Running the protocol instead, the agent looped the test two hundred times, captured failing cases, and noticed failures only when two jobs had the same timestamp. The hypothesis: an unstable sort on ties. Forcing equal timestamps failed every time. The fix was a secondary sort key, plus a regression test.

### Production debugging

Production incidents need extra care. The agent should work from copies of evidence, not live systems. Export sanitized logs and traces into a scratch folder, and strip customer personal data first. If you connect observability tools through MCP, use scoped, read-only, audited credentials. And never let an agent run commands against a production database during an incident. Use a replica or a snapshot.

### Reading unfamiliar errors

Agents are also superb at explaining unfamiliar errors and stack traces. Ask for evidence along with the explanation. Which frame is the last one in our code? What input would take this path? And for library issues, ask the agent to read the installed version's changelog or source, instead of relying on memory, because its memory might describe a different version than the one you actually run.

### Pitfalls

Four pitfalls to avoid. Retries as fixes, which hide real bugs unless the failure is truly external. Shotgun edits, where several changes at once make it impossible to know what worked. Leftover debug logging, which can leak sensitive data. And trusting remembered APIs. Measure your progress with time to root cause, the share of fixes that include a regression test, and bugs that reopen within thirty days.

### Example: blank search page

A simple example. Users report that the search page is blank for some queries. Step one: the agent writes a failing test using the exact query from a user report, one containing an ampersand. Step two: evidence shows the request returns a server error. Step three: hypotheses are unescaped special characters in the query builder, an encoding issue in the front end, or a timeout. Step four: the test with a plain query passes and the one with an ampersand fails, so the escaping hypothesis survives. Step five: a one-line fix to use parameter binding. Step six: the regression test stays forever.

### Scenario: 'random' checkout errors (illustrative)

Now a realistic scenario with illustrative numbers. An online grocery in Dubai has checkout errors roughly twice a day, which support has been blaming on customers' networks. An engineer pairs with an agent, exports a week of sanitized logs, and asks for hypotheses with evidence. The agent spots that every failure happens within a few seconds of a price update job. A reproduction confirms it: the cart total is computed with the old price and validated against the new one. The fix is a price version check. Before, they had spent weeks adding retries. Common mistake: treating frequency as proof of randomness. Rare is not the same as random.

### Deeper: making 'random' deterministic

One level deeper on the Dubai grocery example. The reproduction was the key move. The agent wrote a test that computes a cart total at price version one, then triggers a price update, then validates checkout. It failed deterministically. Once a rare, timing-based bug becomes a deterministic test, the fix, a price version check, takes minutes, and the test guarantees it never quietly returns.

### Watch me do it: the flaky test

Watch me do it with the flaky test from the Islamabad example, using the debug protocol prompt. Step one, reproduce: I give the agent the loop runner script. It runs the test two hundred times and reports eleven failures, then saves the logs of each failing run. Step two, evidence: it compares failing and passing logs and notices every failure contains two jobs with the same created timestamp. Step three, hypotheses: clock resolution in the test environment, an unstable sort when timestamps tie, or transaction isolation letting a job appear twice. For each it proposes a cheap experiment. Step four: it forces equal timestamps in a new test. Fails every time. It forces distinct timestamps. Passes every time. Isolation is ruled out because the job count is always correct. The unstable-sort hypothesis survives. Step five: it proposes adding job ID as a secondary sort key in the query, and waits. I approve. Step six: it implements the one-line change, removes its temporary logging, runs the loop two hundred more times with zero failures, runs the full suite, and keeps the equal-timestamp test as a regression test. Total time, about twenty minutes. Compare that with the first attempt, a retry that would have hidden a real ordering bug from our customers.

### Recap

Recap. Impose the scientific method: reproduce, gather evidence, hypothesize, experiment, fix the cause, and prove it. Use git bisect to find regressions, debug production from sanitized copies, and verify library behavior against the installed version. Your next step: save the debug protocol prompt from the lesson as a skill or snippet, and use it on the next bug your team sees.

## Key takeaways

- No reproduction, no fix: start with a failing test or script.
- Require ranked hypotheses with cheap experiments before any code change.
- Use git bisect with a test command to find regressions automatically.
- Debug production from sanitized copies with scoped read-only access; never against live databases.

## Try it

Save the six-step debug protocol as a reusable prompt or skill, and use it on the next bug, including a regression test.

- [Previous: Tests as guardrails: TDD with agents](https://optimizeall.com/learn/agentic-coding-with-ai/tests-as-guardrails)
- [Next: Code review with AI, and reviewing AI code](https://optimizeall.com/learn/agentic-coding-with-ai/ai-code-review)
- [All lessons of AI-Assisted Software Development: Coding Agents in Practice](https://optimizeall.com/learn/agentic-coding-with-ai)
