AI-Assisted Software Development: Coding Agents in PracticeThe plan-implement-verify workflow · Lesson 8 of 17

Tests as guardrails: TDD with agents

Article · 14 min · 9 min lecture

Video lecture

Tests as guardrails: TDD with agents

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Tests as guardrails

  • TDD with an agent
  • Agent-friendly test suites
  • Stopping agents from cheating

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Tests turn agents from risky to reliable

An agent with no tests is guessing. An agent with a fast, trustworthy test suite has an objective signal it can iterate against for as long as it takes. That makes tests the single most important guardrail in AI-assisted development, and it flips an old habit: test-first development becomes cheap, because the agent writes much of the boilerplate.

Test-driven development with an agent

A reliable pattern, recommended in several vendors' agent best-practice guides:

  1. Write tests first, based on the acceptance criteria. Tell the agent explicitly that you are doing TDD so it does not create mock implementations.
  2. Run them and confirm they fail for the right reason.
  3. Commit the tests.
  4. Implement until tests pass, instructing the agent not to modify the tests.
  5. Review the implementation and whether tests truly capture the requirement.
We are doing test-driven development. Based on the acceptance criteria in issue #231,
write tests in test/tax/ae.test.ts only. Do not write any implementation. Include:
normal cases, the .5 fils rounding boundary, zero and negative amounts, and a non-AE
invoice that must be unchanged. Run the tests and confirm they fail for the expected reason.

Then, in a later step: "Now implement to make these tests pass. Do not modify files under test/."

What makes a test suite agent-friendly

  • Fast. Single-file or single-test runs in seconds. Document the command in AGENTS.md.
  • Deterministic. Flaky tests teach agents to retry or "fix" the wrong thing. Quarantine flakes.
  • Readable failures. Clear assertion messages give the agent (and you) a precise signal.
  • Layered. Unit tests for logic, integration tests for boundaries (DB, queues), a few end-to-end tests for critical journeys, plus snapshot or golden-file tests where output format matters.
  • Beyond examples. Property-based tests (Hypothesis in Python, fast-check in JavaScript) catch edge cases agents and humans both miss.
# Property-based test with Hypothesis: money round-trips must not drift
from decimal import Decimal
from hypothesis import given, strategies as st
from pricing import to_minor_units, from_minor_units

@given(st.decimals(min_value=Decimal("0"), max_value=Decimal("1000000"), places=2))
def test_minor_units_round_trip(amount):
    assert from_minor_units(to_minor_units(amount)) == amount

Guarding the guardrail

Agents under pressure to make tests pass sometimes take shortcuts: weakening assertions, adding skip, special-casing test inputs, or mocking the thing under test. Defenses:

  • Instruction: "Never modify, skip or delete tests to make them pass; if a test seems wrong, stop and explain."
  • Enforcement with hooks: block edits to test files during implementation steps. In Claude Code, for example, a PreToolUse hook can inspect the file path of an edit and reject it:
#!/usr/bin/env bash
# .claude/hooks/protect-tests.sh — reads the tool call JSON on stdin
path=$(jq -r '.tool_input.file_path // empty')
if [[ "$path" == *"/test/"* || "$path" == *".test."* ]] && [[ "${PROTECT_TESTS:-1}" == "1" ]]; then
  echo "Editing tests is blocked during implementation. Explain why the test is wrong instead." >&2
  exit 2   # exit code 2 blocks the tool call and returns the message to the agent
fi
exit 0

Check the current hooks reference for configuration syntax. Other tools offer similar controls or can be wrapped by CI checks.

  • CI checks: fail the build if coverage drops, if tests are newly skipped, or if test files changed in a PR labeled "implementation only".
  • Mutation testing (for example Stryker for JavaScript/TypeScript or mutmut for Python) checks whether tests actually detect injected bugs. Useful for critical modules where you suspect weak AI-written tests.

Worked example: tests that lied

An e-commerce team in Jeddah asked an agent to add tests for a discount engine. Coverage jumped, and every test passed. A mutation-testing run then showed most mutants survived: the tests asserted that functions returned something, not the right amount. The fix was a prompt change ("assert exact values computed by hand for each case; include a table of inputs and expected outputs in the test") and a rule that AI-written tests on money code require a mutation score threshold.

Pitfalls

  • Treating coverage as quality. Coverage says lines ran, not that behavior is checked.
  • Letting the agent write the tests and the code in one step for important logic. You lose the independent check.
  • Slow suites. If tests take 20 minutes, the agent will skip running them or you will.

How to measure success

Track escaped defects on agent-assisted changes, mutation score on critical modules, and the number of PRs where test files changed unexpectedly.

Key takeaways

  • Tests give agents an objective signal; test-first development is now cheap.
  • Write and commit tests first, confirm they fail, then implement without touching tests.
  • Agent-friendly suites are fast, deterministic, readable and layered, with property-based tests for key logic.
  • Protect the signal with instructions, hooks that block test edits, CI checks and mutation testing.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent's PR makes a failing test pass by changing `assertEqual(total, 1999)` to `assertTrue(total > 0)`. What is this?
  2. Which tool best reveals whether AI-written tests actually detect bugs?

Put it into practice

Pick one critical module, run a mutation testing tool against its tests, and strengthen the weakest tests using hand-computed expected values.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.