Skip to content

AI-Assisted Software Development: Coding Agents in Practice · Security, governance and honest measurement · lesson 15 of 17 · 13 min

Measuring AI-assisted productivity honestly

Why honest measurement matters

Claims about AI coding productivity range from "10x" to "it slows us down". Both can be true for different people, tasks and codebases. If you do not measure honestly, you will either overinvest in tools that feel fast or miss real gains. The goal is not to prove AI works; it is to learn where it helps your team and where it does not.

What the evidence says (and does not)

  • Controlled studies vary widely. Early vendor-run experiments on small, well-defined tasks reported large speed-ups. METR's randomized controlled trial with experienced open-source developers on their own mature repositories (early-2025 tools) found tasks took longer on average with AI, while developers believed they were faster; METR's later, larger cohort showed a smaller, statistically uncertain effect and noted selection problems as developers declined to work without AI.
  • Context matters. Familiarity with the codebase, task type, test quality and tool maturity all change outcomes.
  • Self-report is unreliable. Perceived speed-ups and measured outcomes diverge.

Takeaway: do not import someone else's number. Measure your own, carefully.

A balanced metric set

Use several lenses, because any single metric can be gamed.

| Lens | Example metrics | Watch for | |---|---|---| | Delivery (DORA) | Lead time for changes, deployment frequency, change failure rate, time to restore service | Faster lead time with higher change failure rate = debt | | Flow | PR cycle time, review time, PR size, rework rounds | Review time rising as agents generate more code | | Quality | Escaped defects, reverted PRs, incidents, bugs reopened | Quality lagging by weeks | | Cost | Tool spend, tokens per merged change | Runaway agents, low-value usage | | Experience | Developer survey (satisfaction, cognitive load, confidence) | Burnout from constant review |

Frameworks such as DORA's four keys and the SPACE framework (satisfaction, performance, activity, communication, efficiency) help structure this. Avoid lines of code and number of AI suggestions accepted as success metrics; they reward volume, not value.

Designing a fair comparison

  • Baseline first. Capture 6–8 weeks of metrics before broad rollout.
  • Compare like with like. Segment by task type (bug, feature, refactor) and repo.
  • Use cohorts or staggered rollout. Teams adopting at different times give a natural comparison.
  • Track lagging indicators. Defects and incidents surface weeks later.
  • Record mode and tool per PR (a label like ai:delegate) so you can analyze where AI helps.

Hands-on: pull PR metrics with the GitHub CLI

# Merged PRs in the last 30 days with labels, size and timestamps (requires gh auth)
gh pr list --state merged --limit 500 \
  --search "merged:>=$(date -d '30 days ago' +%F)" \
  --json number,title,labels,additions,deletions,createdAt,mergedAt \
  > prs.json
# analyze_prs.py: cycle time and size by AI label
import json, statistics
from datetime import datetime

prs = json.load(open("prs.json"))
def hours(p):
    fmt = "%Y-%m-%dT%H:%M:%SZ"
    return (datetime.strptime(p["mergedAt"], fmt) - datetime.strptime(p["createdAt"], fmt)).total_seconds() / 3600

groups = {}
for p in prs:
    labels = {l["name"] for l in p["labels"]}
    key = next((l for l in labels if l.startswith("ai:")), "ai:none")
    groups.setdefault(key, []).append(p)

for key, items in sorted(groups.items()):
    ct = [hours(p) for p in items]
    size = [p["additions"] + p["deletions"] for p in items]
    print(f"{key:12} n={len(items):3}  median cycle h={statistics.median(ct):6.1f}  median size={statistics.median(size):6.0f}")

Combine with defect data from your issue tracker (bugs linked to PRs) for the quality lens.

Worked example: the review bottleneck

A marketing-tech company in Abu Dhabi saw PR volume rise sharply after adopting agents, and leaders celebrated. The balanced dashboard told a different story: median review time doubled, PR size grew, and change failure rate ticked up. The fix was not fewer agents but better practice: PR size limits, the self-review step, AI first-pass review, and specs with negative criteria. Two months later, throughput stayed high and review time returned to baseline.

Vibe metrics to avoid

  • "Percentage of code written by AI." Interesting, not a goal.
  • "Hours saved" estimated by developers.
  • Leaderboards of individual AI usage. They encourage gaming and erode trust.

How to measure success

A quarterly report that answers three questions with data: where does AI help us most, where does it hurt, and what will we change next quarter?

Video lecture: Measuring AI-assisted productivity honestly

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

  1. Measuring productivity honestly
  2. Analogy: the faster oven
  3. The evidence varies
  4. What to take away
  5. Five lenses
  6. Vanity metrics to avoid
  7. A fair comparison
  8. Hands-on + case
  9. The fix: better practice
  10. Give it time
  11. Try it: a one-hour baseline
  12. Common mistakes
  13. Deeper: a simple size limit
  14. Watch me do it: PR analysis (illustrative)
  15. Recap

Lecture transcript

Measuring productivity honestly

You have probably heard that AI makes developers ten times faster. You may also have heard that it slows them down. Both claims can be true, for different people, tasks and codebases. In this lecture you will learn what the evidence actually says, a balanced set of metrics that cannot easily be gamed, and how to run a fair comparison on your own team.

Analogy: the faster oven

Here is an analogy for measuring productivity. Think about a restaurant that buys a new, faster oven. The chef says dishes come out much quicker. But are customers served sooner? Maybe the bottleneck just moved to plating, or to the one waiter. And are more dishes being sent back? If you only time the oven, you will celebrate while the dining room waits. Measuring AI in software is the same. Generation is the oven. Review, testing and production quality are plating, serving and returns.

The evidence varies

Start with the evidence. Early vendor experiments on small, well-defined tasks reported big speed-ups. Then METR ran a randomized controlled trial with experienced open-source developers working on their own mature repositories, using early twenty twenty-five tools. On average, tasks took longer with AI, while the developers believed they had been faster. A later, larger cohort showed a smaller and statistically uncertain effect, and METR noted a selection problem: many developers no longer wanted to work without AI at all.

What to take away

The lesson is not that AI is slow or fast. It is that context matters enormously. Familiarity with the codebase, the type of task, the quality of tests and the maturity of the tools all change the outcome. And self-reports are unreliable. So do not import someone else's number. Measure your own, carefully.

Five lenses

Use several lenses, because any single metric can be gamed. Delivery, using the DORA four keys: lead time, deployment frequency, change failure rate and time to restore. Flow: pull request cycle time, review time, size and rework. Quality: escaped defects, reverts and incidents. Cost: tool spend and tokens per merged change. And experience: a developer survey on satisfaction, cognitive load and confidence. The SPACE framework is another useful structure.

Vanity metrics to avoid

And avoid vanity metrics. Lines of code and number of suggestions accepted reward volume, not value. Percentage of code written by AI is interesting, but it is not a goal. Hours saved, estimated by developers, is exactly the self-report problem again. And please avoid individual AI usage leaderboards. They encourage gaming and erode trust.

A fair comparison

To compare fairly, capture a baseline of six to eight weeks before broad rollout. Compare like with like, segmenting by task type and repository. Use staggered rollout, where teams adopt at different times, to create a natural comparison. Track lagging indicators like defects, which appear weeks later. And label each pull request with the AI mode used, so you can see where AI helps and where it does not.

Hands-on + case

The lesson includes a small script that pulls merged pull requests with the GitHub command line tool and reports median cycle time and size by AI label. It takes minutes to run. Then a real example. A marketing-tech company in Abu Dhabi saw pull request volume jump after adopting agents, and leaders celebrated. The balanced dashboard told another story: review time doubled, pull requests got bigger, and change failure rate crept up.

The fix: better practice

They did not remove the agents. They improved the practice. They set pull request size limits, added the fresh-eyes self-review step, put AI first-pass review in CI, and required negative criteria in specs. Two months later, throughput stayed high and review time returned to its baseline. That is the point of measuring honestly. It tells you what to fix, not just whether to cheer.

Give it time

One more practical point: quality signals lag. A change merged today may cause an incident next month. So give any comparison enough time, at least one or two release cycles, before drawing conclusions, and link bugs back to the pull requests that introduced them. Share results openly with the team, including the uncomfortable ones. People trust a measurement program that reports where AI hurt as readily as where it helped.

Try it: a one-hour baseline

A simple example you can compute this week. Take last month's merged pull requests. Label each as AI-assisted or not, using the labels from this lesson going forward and your best memory for the past. Compute the median time from opening to merge for each group, and the median size. If AI-assisted pull requests are much bigger and take longer to review, you have learned something important without any fancy tooling. If they are similar in size and faster, you have early evidence of real gains. Either way, you now have a baseline.

Common mistakes

Common mistakes, and they are tempting. Announcing a productivity number to leadership after two weeks, before quality signals have had time to appear. Comparing a senior team using AI with a junior team without it, which measures seniority. Rewarding individuals for AI usage, which produces gaming, not value. And quietly dropping the metrics that look bad. The healthiest teams publish what they learn, good and bad, and treat the numbers as a guide for improving practice, not as a scoreboard.

Deeper: a simple size limit

One level deeper on the Abu Dhabi case. The size limit they chose was simple: pull requests over four hundred changed lines need a written reason and a second reviewer. Within a month, most agent work arrived in smaller slices, because developers started asking agents to stop after each plan step. The metric changed the habit, which is exactly what good metrics should do.

Watch me do it: PR analysis (illustrative)

Watch me do it. I'm running the pull request analysis from the lesson on last month's work. First, the GitHub command line query: merged pull requests from the last thirty days with labels, additions, deletions and timestamps, saved to a JSON file. Five hundred rows at most; we have one hundred and forty. Second, I run the Python script. It groups by the AI label and prints three rows. No AI: seventy-two pull requests, median cycle time about nineteen hours, median size about ninety lines. Pair mode: forty-one pull requests, median about seventeen hours, size about one hundred and ten. Delegate mode: twenty-seven pull requests, median about twenty-six hours, size about two hundred and forty. Those numbers are illustrative, but look at the pattern. Delegated work is bigger and waits longer. Third, I dig in: I sort the delegate group by size and open the five largest. Three came from specs without negative criteria and touched unrelated files. Fourth, I check quality: I join bugs from the issue tracker to pull request numbers and find two escaped defects in the delegate group, zero in pair. So my one-page report says: pair mode is helping, delegate mode needs better specs and size limits, and we will measure again next month after adding both.

Recap

Recap. The evidence varies, so measure your own team. Use delivery, flow, quality, cost and experience together. Avoid vanity metrics. Baseline first, segment, stagger and label. Your next step: add AI mode labels to your pull requests this week, run the script from the lesson after a month, and write a one-page report answering where AI helps, where it hurts, and what you will change.

Key takeaways

  • Published results vary widely; measure your own team rather than importing a number.
  • Use balanced lenses: delivery (DORA), flow, quality, cost and developer experience.
  • Avoid vanity metrics such as lines of code, suggestions accepted or self-estimated hours saved.
  • Baseline first, segment by task, stagger rollout, track lagging quality indicators, and label PRs by AI mode.

Try it

Add AI mode labels to PRs for a month, run the analysis script from the lesson, and write a one-page 'where it helps, where it hurts' report.