AI-Assisted Software Development: Coding Agents in PracticeThe coding-agent landscape and how agents work · Lesson 3 of 17

Choosing and setting up tools with a fair trial

Article · 12 min · 9 min lecture

Video lecture

Choosing and setting up tools with a fair trial

13 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 13

Choosing and setting up your tools

  • Criteria that matter
  • A fair two-week trial
  • Day-one setup checklist

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Choosing is a trial, not a debate

Teams waste weeks arguing about which assistant is "best". The productive question is narrower: which tool, configured how, helps this team ship reviewed, tested changes on this codebase? Answer it with a structured two-week trial.

Evaluation criteria that actually matter

CriterionQuestions to ask
Fit with workflowTerminal, IDE, or cloud? Does it work with your editor and your Git host?
Context handlingDoes it read AGENTS.md or its own instruction file? Can it handle your repo size? Does it support MCP for your issue tracker and docs?
ControlsSandboxing, approval modes, allow/deny lists, hooks, audit logs, admin policies
Data and privacyRetention, training use, region, zero-data-retention options, SSO, contractual terms for your plan
Cost modelPer-seat, usage-based, or both? What happens at heavy usage? Can you set budgets?
CI and automationHeadless mode, official GitHub Action or equivalent, API access
Model choiceLocked to one vendor's models or configurable?

Weight the criteria for your context. An agency handling client code in the UAE and UK will weight data terms and auditability heavily. A solo creator building a Shopify app will weight cost and speed of setup.

Designing a fair trial

  1. Pick 15–30 real tickets from your backlog across sizes and types: bug fixes, small features, tests, refactors, docs.
  2. Define "done" identically for every tool: merged PR, CI green, reviewer approved without major rework.
  3. Record per ticket: tool, mode (pair/delegate/pipeline), wall-clock time, reviewer minutes, number of revision rounds, defects found later.
  4. Rotate developers across tools so one expert does not skew the result.
  5. Decide in advance what result would make you adopt, extend the trial, or reject.

Beware the novelty effect and self-reported speed. A randomized study by METR on experienced open-source developers in early 2025 found participants took longer with AI tools on average, while believing they were faster; METR's later work with a larger cohort showed a smaller, statistically uncertain effect. The lesson is not "AI is slow" but "measure, don't guess".

Setting up for success on day one

Whichever tool you choose, the same setup checklist applies:

  • Instruction file at the repo root (AGENTS.md, plus CLAUDE.md or tool-specific files where needed) with build, test and style commands. Module 2 covers this in depth.
  • Permission baseline committed to the repo.
  • Secrets hygiene: .env files ignored by git and denied to the agent; use a secrets manager for real credentials.
  • Fast, reliable tests runnable with one command. Agents amplify whatever feedback loop you give them.
  • Branch protection on main: required reviews, required status checks, no direct pushes.
  • Budget alerts on usage-based plans and API keys.

Hands-on: install and smoke-test two agents

Pick two tools and run the same smoke test in a throwaway branch. Commands change; confirm against each vendor's install docs.

# Example: Claude Code (see Anthropic's install docs for your OS)
npm install -g @anthropic-ai/claude-code
cd your-repo && git switch -c trial/claude
claude   # interactive; try: "Explain how to run the tests here, then run them."

# Example: Codex CLI (see OpenAI's docs)
npm install -g @openai/codex
cd your-repo && git switch -c trial/codex
codex    # same prompt

Then give both the identical task:

Read AGENTS.md. Find one function in src/ without unit tests. Write focused tests for it
covering normal, edge and error cases. Run the test suite. Do not modify src/.
Summarize what you tested and anything suspicious you noticed.

Compare: did each agent follow the constraint (no changes to src/)? Did the tests actually fail when you deliberately break the function? How long did your review take?

Worked example: an agency's decision

A 12-person agency in Dubai builds WordPress and Next.js sites for clients in the UAE and KSA. Their trial scored three tools on 20 tickets. The winner was not the one with the most impressive demo but the one whose admin controls let them block specific client repositories from cloud agents, enforce SSO, and export audit logs for client security questionnaires. They still let developers use a second tool locally for exploration, with a written policy on which repositories are allowed.

Pitfalls

  • Trial on toy projects. Greenfield demos flatter every tool.
  • Ignoring reviewer time. Code that takes 5 minutes to generate and 60 to review is not a win.
  • One tool for everything. Many teams settle on one primary agent plus a CI reviewer.

How to measure success

A decision memo backed by your ticket log: adoption decision, configuration baseline, allowed repositories, and the metrics you will keep tracking (Module 5).

Key takeaways

  • Choose with a structured trial on real tickets, not demos or leaderboards.
  • Weight criteria for your context: controls and data terms matter as much as capability.
  • Log reviewer time, rework and later defects; self-reported speed is unreliable.
  • Apply the same day-one setup: instruction file, permission baseline, secrets denied, fast tests, branch protection, budget alerts.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which metric is most often missed when teams compare coding agents?
  2. What did METR's early-2025 randomized study of experienced open-source developers highlight?

Put it into practice

Run the two-agent smoke test from this lesson on a throwaway branch and record constraint-following, test quality and your review time.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.