AI-Assisted Software Development: Coding Agents in PracticeContext engineering for codebases · Lesson 6 of 17
Writing tasks agents can finish
Video lecture
Writing tasks agents can finish
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Writing tasks agents can finish
Give an agent a vague ticket and it will confidently solve a different problem. Give it a sharp one and it can open a mergeable pull request while you are in a meeting. In this lecture you will learn the anatomy of an agent-ready task, how to size work so agents can finish it, and a reusable issue template for your team.
0:27 Analogy: the creative brief
Why does this matter so much? Because an agent is like a brilliant new colleague who arrived this morning, cannot see your chat history, and will not interrupt you with questions. Think about briefing a freelance designer in Lahore for a logo. If you say make it pop, you get something that pops in a direction you did not want. If you give the brand colors, the audience, three examples you like and one you hate, you get something usable in the first round. A task spec is that creative brief for code.
1:07 The spec is the product
With agents, the task spec is the product. The output can only be as good as the description. Teams that track this, including published research on Copilot's issue-to-pull-request agent, see the same pattern: issues with clear scope, acceptance criteria and pointers to relevant code are much more likely to produce mergeable changes. Writing a good task is no longer admin. It is engineering.
1:34 Six parts
An agent-ready task has six parts. The goal says what and why. The context lists the relevant files, similar past pull requests and docs, which saves exploration. The acceptance criteria define done in checkable terms. The constraints fence the blast radius, like no schema changes or no new dependencies. Verification gives the exact commands that prove success. And out of scope stops the agent from being helpfully wrong.
2:04 Worked example: UAE VAT line
Look at the example in the lesson. Customers in the U A E should see VAT at five percent as a separate invoice line. The context points to the tax module and a previous pull request that added a similar line for Pakistani invoices. Follow that pattern is one of the most powerful phrases you can give an agent. Acceptance criteria include a rounding test at the half-fils boundary, and a statement that non-Emirati invoices must not change. That negative criterion is the one people forget.
2:41 Sizing
Now sizing. Agents succeed most often on tasks a competent developer could finish in under a day and review in minutes. Warning signs of an oversized task: acceptance criteria spanning several subsystems, not knowing which files will change, or the words and also in the spec. Split along seams. Data model first, then API, then UI. Or one package at a time. Every slice should leave the system working and tested.
3:12 Handling ambiguity
Agents rarely ask clarifying questions unless you tell them to. So tell them. Add a line asking the agent to stop and ask if any criterion is ambiguous or conflicts with existing code, and to list its assumptions at the top of its plan. In delegate mode, where it cannot ask you in real time, require the assumptions in the pull request description so the reviewer can check them first.
3:42 Before → after
Here is a before and after. Before: make the export faster. After: the CSV export of fifty thousand invoices times out for large clients in Karachi; the handler loads everything into memory; the database driver supports streaming. Done means memory stays flat in the load test and output stays byte-identical for the fixture file. Same endpoint, same headers, Excel out of scope. Five minutes of writing turned an open-ended hunt into a checkable task.
4:14 More traps
One more trap: specifying the implementation instead of the outcome. If you dictate every line, you have done the work yourself and the agent adds little. Give it the goal, the constraints and the checks, and let it choose the approach, unless the approach genuinely matters, for example for performance or compliance reasons. Also watch for hidden context. If the key decision was made in a chat thread the agent cannot see, it does not exist for the agent. Copy the decision into the issue.
4:51 Drill: 'Add dark mode' → agent-ready
Let's practice on a simple one. The vague ticket says: add dark mode. Take ten seconds and ask what is missing. Here is a sharper version. Goal: users can switch between light and dark themes, and the choice persists. Context: theme tokens live in the styles tokens file; follow the pattern from the high-contrast pull request. Acceptance: a toggle in settings, preference saved per user, all pages pass the contrast checker. Must not change: existing light theme screenshots. Out of scope: email templates. Notice the negative criterion again. It protects the ninety percent of the app you did not want touched.
5:35 Scenario: 20 tickets, two weeks (illustrative)
And a realistic scenario with illustrative numbers. An e-commerce team in Jeddah assigns twenty tickets to a cloud agent over two weeks. Ten are written with the six-part template and ten are their usual one-liners. Nine of the template tickets merge after a single review round. Only four of the one-liners do, and three of those one-liners come back as pull requests that solved the wrong problem entirely. Writing the templates took about five minutes each. The team's conclusion: the template is the cheapest productivity tool they own.
6:13 Deeper: splitting dark mode
One level deeper on sizing. Suppose the dark mode ticket also asked for email templates and a new settings page. Split it along seams. Slice one: theme tokens and the toggle, with screenshot tests. Slice two: persist the preference per user. Slice three: email templates, separately, because they use a different rendering system. Each slice ships green on its own, and each one can be reviewed in minutes.
6:43 Watch me do it: vague → agent-ready
Watch me do it. I'm turning the vague ticket, make the export faster, into the agent-ready template, section by section. Goal: the CSV export of fifty thousand invoices times out after thirty seconds for large clients in Karachi; exports must complete. Context: I open the handler and confirm it loads every row into memory; I note the database driver already supports streaming, and I link the stream helper file. Acceptance criteria, written as checkboxes: rows are streamed; memory stays flat in the load test, and I give the exact command; output is byte-identical to today's export for the small fixture file. Constraints: same endpoint, same headers, no new dependencies. Verification: npm test, npm run check, and the load test command, with results pasted in the pull request. Out of scope: Excel export. Assumptions the agent must list: for example, whether a slow client connection should hold a database transaction open. Now I do the thirty-second review from the lesson. Could a new colleague start without asking a question? Yes. Can every criterion be checked? Yes, all three have commands or a byte comparison. Is there a negative criterion? Yes, same endpoint and headers, and identical output. Is anything hidden in a chat thread? The decision to leave Excel out was made in a meeting, and it is now written down. Label it agent-ready and assign it.
8:21 Recap
Recap. The spec is the product. Use the six parts, size tasks to under a day, include negative criteria, and make the agent surface assumptions. Your next step: add the agent task issue template from the lesson to your repository, and only allow issues that use it to be assigned to cloud agents. Then track first-attempt merge rates with and without the template.
The task spec is the product
With coding agents, the quality of the output is bounded by the quality of the task description. A vague ticket produces a plausible-looking change that solves a different problem. Research on Copilot's issue-to-PR agent and plenty of team experience point the same way: issues with clear scope, acceptance criteria and pointers to relevant code are far more likely to produce mergeable pull requests.
Anatomy of an agent-ready task
## Goal
Customers in the UAE should see VAT (5%) as a separate line on PDF invoices.
## Context
- Tax rules: src/domain/tax/ae.ts (rate already defined as AE_VAT_RATE)
- PDF rendering: src/infra/pdf/invoiceTemplate.tsx
- Similar prior change: PR #412 added GST line for PK invoices — follow that pattern.
## Acceptance criteria
- [ ] Invoices with country=AE show "VAT 5%" line with amount in fils → AED formatting.
- [ ] Non-AE invoices unchanged (snapshot tests still pass).
- [ ] Unit test for ae.ts rounding at .5 fils boundary.
- [ ] PDF snapshot test for an AE invoice.
## Constraints
- Do not change the public API or database schema.
- No new dependencies.
## Verification
Run: npm test && npm run check. Attach the generated sample PDF path in the PR description.
## Out of scope
Arabic-language PDF layout (tracked in INV-233).Each section does a job: Goal tells the agent why; Context saves exploration; Acceptance criteria define done in checkable terms; Constraints fence the blast radius; Verification tells it how to prove success; Out of scope stops helpful over-reach.
Sizing tasks
Agents succeed most often on tasks that a competent developer could finish in under a day and review in minutes. Signs a task is too big:
- The acceptance criteria span more than one subsystem.
- You cannot say which files will change.
- The spec contains "and also".
Split along seams: data model change, then API, then UI; or one package at a time. Each slice should leave the system working and tested.
Writing for ambiguity
Agents rarely ask clarifying questions unless told to. Add an explicit instruction:
If any acceptance criterion is ambiguous or conflicts with existing code, stop and ask
me before implementing. List your assumptions at the top of your plan.In delegate mode, where the agent cannot ask you in real time, require it to list assumptions in the PR description so reviewers can check them.
Worked example: from vague to agent-ready
Before: "Make the export faster."
After:
- Goal: CSV export of 50k invoices currently times out (>30s) for large Karachi-based clients.
- Context:
src/api/handlers/exportCsv.tsloads all rows into memory; the DB driver supports streaming (seesrc/infra/db/stream.ts). - Acceptance: export streams rows; memory stays flat in the provided load test (
npm run test:load -- export); output byte-identical to current export for fixturetest/fixtures/export-small.json. - Constraints: same endpoint and response headers.
- Out of scope: Excel export.
The rewrite took five minutes and turned an open-ended performance hunt into a checkable engineering task.
Hands-on: an issue template for agent tasks
Save as .github/ISSUE_TEMPLATE/agent-task.md:
---
name: Agent-ready task
about: A bounded task suitable for a coding agent
labels: agent-ready
---
## Goal
## Context (files, prior PRs, docs)
## Acceptance criteria
- [ ]
## Constraints (API, schema, dependencies, performance)
## Verification (exact commands)
## Out of scope
## Assumptions the agent must list in the PRUse the agent-ready label as a gate: only issues with this template filled in may be assigned to cloud agents.
Pitfalls
- Specifying the implementation instead of the outcome. Give constraints, not line-by-line instructions, unless the approach genuinely matters.
- Hidden context. Decisions from a Slack thread the agent cannot see.
- No negative criteria. Say what must not change.
Reviewing the spec before assigning
Before you hand a task to a cloud agent, run a thirty-second review against this checklist: Could a new colleague start without asking a question? Can every acceptance criterion be checked by a command or a clearly observable behavior? Is there at least one negative criterion? Are the files or modules named? Is anything the agent needs locked away in a chat thread or a meeting? If any answer is no, fix the spec first. It is far cheaper to spend two minutes improving the issue than twenty minutes reviewing and rejecting a pull request built on a misunderstanding. Some teams also ask the agent itself to critique the spec before starting: "List anything in this issue that is ambiguous or missing." That single step catches many gaps.
How to measure success
Track first-attempt merge rate for agent PRs, split by whether the issue used the template. If the template is working, the difference will show within a few weeks.
Key takeaways
- The task spec bounds output quality: goal, context, acceptance criteria, constraints, verification, out of scope.
- Size tasks to under a day of work and minutes of review; split along seams.
- Include negative criteria (what must not change) and pointers to similar prior PRs.
- Require agents to ask or list assumptions when criteria are ambiguous.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Rewrite one vague ticket from your backlog into the six-part agent-ready format and add the issue template to your repository.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.