Skip to content

AI-Assisted Software Development: Coding Agents in Practice · From vibe coding to shipping: capstone · lesson 17 of 17 · 20 min

Capstone: ship a feature end to end with an agent

The brief

You will ship one real feature end to end with a coding agent, applying everything in this course. Choose a codebase you control (a work repository with permission, a side project, or the reference project below). The deliverable is a merged pull request plus a short written retrospective.

Reference project if you need one: a small API for a creator storefront (Python FastAPI or Node/Express, your choice) with products and orders. The feature: discount codes with percentage or fixed amounts, expiry dates, per-customer usage limits, and currency-safe calculations for PKR, AED, SAR, GBP and USD.

Step 1: set up the repository (Modules 1–2)

  • Create or update AGENTS.md (map, commands, conventions, done, boundaries, gotchas) and a tool-specific pointer file if needed.
  • Commit a permission baseline; deny .env*.
  • Add secret scanning (pre-commit) and branch protection.
  • Confirm tests run with one command.

Step 2: write the agent-ready spec (Module 2)

## Goal
Customers can apply a discount code at checkout.
## Context
- Orders: src/orders/ ; money helpers: src/money.py (integer minor units)
## Acceptance criteria
- [ ] Codes: percent (1–100) or fixed amount in the order currency; optional expiry (UTC).
- [ ] Per-customer usage limit enforced atomically (no double use under concurrent requests).
- [ ] Discount never makes total negative; rounding half-up in minor units.
- [ ] Expired, unknown or over-limit codes return 422 with a clear error code.
- [ ] Orders without codes behave exactly as before (existing tests unchanged and passing).
## Constraints
- No new dependencies. No schema changes outside a new `discount_codes` and `discount_redemptions` table.
## Verification
- pytest -q (or npm test); new tests cover every criterion; include a concurrency test.
## Out of scope
- Admin UI for creating codes (seed via fixture).

Step 3: explore and plan (Module 3)

Explore only. Read AGENTS.md and the files in Context. Summarize the order flow and where
a discount should apply, with file:line citations. Then propose a plan: failing tests first,
then schema, then domain logic, then API. List risks (especially concurrency) and assumptions.
Stop after the plan.

Review and edit the plan. Save it as PLAN.md on the branch.

Step 4: tests first, then implementation (Module 3)

  1. Ask the agent to write tests for every acceptance criterion, including a concurrency test for usage limits (e.g., two simultaneous redemptions of a single-use code; exactly one succeeds). Confirm they fail. Commit.
  2. Enable test protection (instruction plus hook or CI check).
  3. Implement plan steps one at a time, committing after each verified step.

A property-based test is a strong addition:

from hypothesis import given, strategies as st
from src.discounts import apply_discount

@given(total=st.integers(min_value=0, max_value=10_000_000),
       percent=st.integers(min_value=1, max_value=100))
def test_discount_never_negative_and_never_increases(total, percent):
    result = apply_discount(total_minor=total, kind="percent", value=percent)
    assert 0 <= result <= total

Step 5: verify and review (Modules 3–4)

  • Full suite, type checks, lint: paste the evidence.
  • Run the app and exercise the endpoint manually (valid, expired, over-limit codes).
  • Run the fresh-eyes self-review prompt in a new session.
  • Check the diff yourself with the human review checklist: intent, tests first, plausible fakes, dependencies, security (can a customer use another customer's code? can the limit be bypassed?), duplication.

Step 6: ship through CI (Module 4)

  • Open the PR with a description: what, why, how verified, assumptions, AI tool and mode used.
  • Let the AI review workflow comment; address high-severity findings.
  • Human approval, then merge.

Step 7: retrospective (Module 5)

Write one page:

## Capstone retrospective
- Time: spec __ min, planning __, implementation __, review __, total __
- Agent sessions: __ ; restarts due to lost context: __
- Where the agent helped most / least:
- Bugs caught by: tests __, self-review __, AI review __, human review __
- Instruction-file lines added because of mistakes:
- What I would do differently:

Assessment rubric

| Criterion | Excellent | |---|---| | Context | AGENTS.md is specific and verified; spec has negative criteria and verification commands | | Process | Plan reviewed and saved; tests committed before implementation; small commits | | Quality | All criteria tested including concurrency and property tests; no weakened tests | | Security | No secrets; authorization and limit bypass considered; dependencies unchanged | | Delivery | PR description complete; AI review addressed; human-approved merge | | Reflection | Retrospective uses real numbers and names concrete improvements |

Common capstone pitfalls

  • Skipping the concurrency test (the most common real bug in discount systems).
  • Letting the agent write tests and implementation together.
  • A single huge commit.
  • A retrospective without numbers.

What next

Repeat the loop on a second feature in delegate mode: write the spec, assign it to a cloud agent, and compare your review experience with this pair-mode run.

Video lecture: Capstone: ship a feature end to end with an agent

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

  1. Capstone: ship a feature end to end
  2. Why a capstone?
  3. The feature: discount codes
  4. Steps 1–2
  5. Step 3: explore and plan
  6. Step 4: tests first
  7. Step 5: verify and review
  8. Steps 6–7
  9. Rubric
  10. Example: the double redemption
  11. Retro example (illustrative)
  12. Deeper: enforcing the negative criterion
  13. Watch me do it: the concurrency test
  14. Recap

Lecture transcript

Capstone: ship a feature end to end

This is where everything comes together. In this capstone you will ship one real feature end to end with a coding agent, from repository setup to a merged, human-approved pull request, and then write an honest retrospective. By the end you will have a working feature, a reusable workflow, and real numbers about how agents perform on your code.

Why a capstone?

Before we start, why a capstone at all? Because knowing each technique separately is like knowing the rules of cricket without ever playing a match. The real skill is sequencing: when to write the spec, when to stop and review the plan, when to lock the tests, when to switch from agent to human judgment. Playing one full match, start to finish, on a real feature, is what turns the separate lessons into a workflow you can repeat on Monday morning.

The feature: discount codes

Use a codebase you control, or the reference project: a small API for a creator storefront. The feature is discount codes, with percentage or fixed amounts, expiry dates, per-customer usage limits, and currency-safe calculations for rupees, dirhams, riyals, pounds and dollars. It sounds simple. It hides exactly the kinds of bugs agents and humans miss: rounding, negative totals, and two people redeeming a single-use code at the same moment.

Steps 1–2

Step one is setup. Write your instruction file with map, commands, conventions, definition of done, boundaries and gotchas. Commit a permission baseline that denies env files. Add secret scanning and branch protection. And confirm the tests run with a single command. Step two is the spec. The lesson gives you a full agent-ready spec with acceptance criteria, including a negative criterion that orders without codes behave exactly as before, constraints, verification commands and an out of scope section.

Step 3: explore and plan

Step three: explore and plan. Use the read-only exploration prompt, insist on file and line citations, and ask for a plan that starts with failing tests, lists risks, especially concurrency, and states assumptions. Then do your most valuable review of the whole project. Edit the plan and save it as PLAN dot M D on the branch.

Step 4: tests first

Step four: tests first. Have the agent write tests for every criterion, including a concurrency test: two simultaneous redemptions of a single-use code, and exactly one must succeed. Add the property-based test from the lesson, which checks that a discount never makes a total negative and never increases it, across thousands of generated inputs. Confirm the tests fail, commit them, and switch on test protection. Then implement one plan step at a time, committing after each verified step.

Step 5: verify and review

Step five: verify and review. Paste evidence from the full suite, type checks and lint. Run the app and try valid, expired and over-limit codes by hand. Run the fresh-eyes self-review in a new session. Then apply the human checklist, with special attention to security. Can a customer use another customer's code? Can the usage limit be bypassed with parallel requests or by changing case in the code string?

Steps 6–7

Step six: ship through CI. Open the pull request with a complete description: what changed, why, how you verified it, assumptions, and which AI tool and mode you used. Let the AI review workflow comment, address high-severity findings, then get human approval and merge. Step seven: the retrospective. Record time spent on spec, planning, implementation and review. Count restarts. Note which layer caught each bug: tests, self-review, AI review or human review.

Rubric

You will be assessed on six criteria. Context: a specific, verified instruction file and a spec with negative criteria. Process: a reviewed plan, tests committed first, small commits. Quality: every criterion tested, including concurrency and properties, with no weakened tests. Security: no secrets, bypasses considered, no surprise dependencies. Delivery: complete description and approved merge. And reflection: a retrospective with real numbers and concrete improvements. The most common miss, by far, is skipping the concurrency test.

Example: the double redemption

Here is a simple example of the kind of bug the capstone is designed to surface. Two customers in different browsers apply the same single-use code, SAVE10, within the same second. A naive implementation checks whether the code has been used, sees no, and then records the use. Both requests pass the check before either records anything. Result: two redemptions of a single-use code. The fix is an atomic operation, like a unique constraint or a conditional update, plus the concurrency test that proves it. If your first test run does not catch this, your tests are not done.

Retro example (illustrative)

And a realistic picture of what a good retrospective looks like, with illustrative numbers. Spec, twenty minutes. Planning and plan review, twenty-five. Tests, forty. Implementation across five small commits, seventy. Verification and self-review, thirty. Human review, twenty-five. Bugs caught: tests caught three, self-review caught one, the AI reviewer caught one naming issue, and the human reviewer caught the case-sensitivity bypass. Instruction file lines added: two. That page of numbers is worth more than any opinion about whether agents are good.

Deeper: enforcing the negative criterion

Let's deepen the spec step. The negative criterion, orders without codes behave exactly as before, is enforced by leaving every existing checkout test untouched and passing. In the capstone, ask the agent at the end to show a git diff of the existing test folder. It should be empty. If it is not, you have found the most important review comment before anyone else does.

Watch me do it: the concurrency test

Watch me do it: the concurrency test for single-use discount codes, the part most people skip. First, the test. I create a code, SAVE10, with a usage limit of one for customer A. Then I start two requests at the same moment, both trying to redeem it, using a thread pool so they genuinely overlap. The assertion: exactly one returns success, the other returns the over-limit error, and the redemptions table has exactly one row. I run it against the agent's first implementation. It fails about one run in three: two successes, two rows. Now the diagnosis. The code reads the redemption count, checks it is below the limit, then inserts a row. Two requests can both read zero before either inserts. Next, the fix. I ask the agent for an atomic approach and it proposes a unique constraint on code and customer for single-use codes, plus a conditional update on a remaining-uses counter for multi-use codes, so the database rejects the second attempt. Then I run the concurrency test fifty times in a loop. Zero failures. Finally, I add the case-sensitivity test from the review: save10 and SAVE10 must count as the same code, so codes are normalized to uppercase at both creation and lookup. Two small tests, two real bugs that would have cost the business money.

Recap

Recap. Setup, spec, plan, tests first, step-by-step implementation, verification, review, CI and a retrospective. That is the whole discipline of AI-assisted engineering in one feature. Your next step after the capstone: repeat the loop on a second feature in delegate mode. Write the spec, assign it to a cloud agent, and compare the review experience with this pair-mode run. That comparison will teach you more about your team's future workflow than any benchmark.

Key takeaways

  • Set up context first: AGENTS.md, permissions, secret scanning, branch protection, one-command tests.
  • Write an agent-ready spec with negative criteria and verification commands.
  • Plan, commit tests first (including concurrency and property tests), then implement in small verified steps.
  • Verify with evidence, self-review with fresh eyes, ship through CI with human approval, and retrospect with real numbers.

Try it

Complete the capstone: ship the discount-code feature (or your own) via a human-approved PR and write the one-page retrospective with real numbers.