AI-Assisted Software Development: Coding Agents in PracticeThe plan-implement-verify workflow · Lesson 9 of 17

Debugging with agents: a scientific protocol

Article · 13 min · 8 min lecture

Video lecture

Debugging with agents: a scientific protocol

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Debugging with agents

  • A six-step protocol
  • git bisect with an agent
  • Safe production debugging

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Debugging is where agents shine, if you run it like science

Debugging suits agents well: they can read stack traces, grep through code, add logging, and rerun a reproduction dozens of times without getting tired. They are also prone to a classic failure: changing code until the symptom disappears without understanding the cause. Your job is to impose the scientific method.

The debugging protocol

  1. Reproduce. Get a failing test or a script that reliably shows the bug. No reproduction, no fix.
  2. Gather evidence. Logs, traces, inputs, recent changes (git log -p, git bisect).
  3. Hypothesize. List candidate causes ranked by likelihood, each with a way to confirm or refute it.
  4. Test hypotheses. Add targeted logging or assertions; run the reproduction; eliminate causes.
  5. Fix the cause, not the symptom, with the smallest change.
  6. Prove it. The reproduction now passes; the full suite passes; add a regression test.

Encode this as a reusable prompt or skill:

Debug protocol. Do not change application code until step 5.
1) Reproduce the bug as a failing automated test and show me it fails.
2) Collect evidence: relevant logs, inputs, and recent commits touching these files.
3) List 3 hypotheses ranked by likelihood, each with a cheap experiment to test it.
4) Run the experiments (temporary logging is fine) and report which hypothesis survives.
5) Propose the minimal fix for the root cause and wait for my approval.
6) After approval: implement, remove temporary logging, run the full suite, keep the test.

When "it worked last week", git bisect finds the commit that introduced a bug by binary search. Agents can run it automatically when you give them a test command:

git bisect start
git bisect bad HEAD
git bisect good v2.14.0
git bisect run npx vitest run test/export.test.ts
git bisect reset

Ask the agent to read the offending commit and explain why it broke the behavior before fixing.

Debugging production issues safely

For production incidents, the agent should work from copies of evidence, not live systems:

  • Export sanitized logs or traces into the repo's scratch folder; strip customer PII first.
  • Give read-only access to observability tools through MCP only if the credentials are scoped and audited.
  • Never let an agent run commands against production databases during an incident. Use a replica or snapshot.

Worked example: the intermittent failure

A logistics startup in Islamabad saw a test that failed roughly one run in twenty in CI. A developer asked an agent to "fix the flaky test"; its first attempt added a retry, which hid the problem. Running the protocol instead:

  • Reproduction: the agent ran the test 200 times in a loop and captured failing seeds.
  • Evidence: failures occurred only when two jobs shared a timestamp.
  • Hypotheses: (a) clock resolution, (b) unstable sort on equal timestamps, (c) database isolation.
  • Experiment: forcing equal timestamps failed every time; the sort was not stable on ties.
  • Fix: add a secondary sort key (job ID) in the query. Regression test pinned equal timestamps.

The retry would have shipped the real bug to customers, whose shipments occasionally appeared out of order.

Reading unfamiliar errors and stacks

Agents are excellent at explaining unfamiliar errors, framework internals and cryptic stack traces. Ask for the explanation and the evidence: "Which line in our code is the last frame we control? What input would cause this path?" For library bugs, ask the agent to check the installed version's changelog or source in node_modules or site-packages rather than relying on memory, which may describe a different version.

Hands-on: a loop runner for flaky tests

#!/usr/bin/env bash
# run-many.sh <n> <test command...>: run a test repeatedly and count failures
n=$1; shift; fails=0
for i in $(seq 1 "$n"); do
  if ! "$@" > /tmp/run-$i.log 2>&1; then fails=$((fails+1)); echo "fail on run $i"; fi
done
echo "$fails failures out of $n runs"

Give this to the agent as a tool for reproducing intermittent bugs.

Pitfalls

  • Retry as a fix. Retries hide real bugs unless the failure is truly external and transient.
  • Shotgun edits. Several simultaneous changes make it impossible to know what fixed it.
  • Leftover debug logging containing sensitive data. Make removal part of the protocol.
  • Trusting remembered APIs. Versions differ; read the installed code.

How to measure success

Track mean time to root cause, the share of bug fixes that include a regression test, and bugs that reopen within 30 days.

Key takeaways

  • No reproduction, no fix: start with a failing test or script.
  • Require ranked hypotheses with cheap experiments before any code change.
  • Use git bisect with a test command to find regressions automatically.
  • Debug production from sanitized copies with scoped read-only access; never against live databases.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent 'fixes' an intermittent test by adding three retries. What should you do?
  2. During an incident, which access is appropriate for a coding agent?

Put it into practice

Save the six-step debug protocol as a reusable prompt or skill, and use it on the next bug, including a regression test.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.