---
title: "Code review with AI, and reviewing AI code"
description: "Two directions of AI code review AI review cuts both ways: - AI reviewing human (or agent) code. Automated first-pass review on pull requests: GitHub…"
url: https://optimizeall.com/learn/agentic-coding-with-ai/ai-code-review
updated: 2026-10-05
---

AI-Assisted Software Development: Coding Agents in Practice · Code review, migrations and CI integration · lesson 10 of 17 · 13 min

# Code review with AI, and reviewing AI code

## Two directions of AI code review

AI review cuts both ways:

- **AI reviewing human (or agent) code.** Automated first-pass review on pull requests: GitHub Copilot code review, Cursor's Bugbot, Claude Code or Codex running in CI, and dedicated review products. Good at catching slips humans skim past: unhandled errors, missing null checks, inconsistent naming, obvious security smells, missing tests.
- **Humans reviewing agent code.** As agents write more code, human review becomes the bottleneck and the last line of defense. It needs a different technique from reviewing a colleague's code.

## What AI review is good at, and what it is not

| Strong | Weak |
|---|---|
| Local bugs visible in the diff | Whether the change solves the right problem |
| Style and consistency with nearby code | Architectural fit across services |
| Missing error handling and input validation | Business rules not written down anywhere |
| Common security patterns (injection, secrets) | Subtle concurrency and distributed-systems issues |
| Summarizing large diffs | Knowing which trade-offs your team has accepted |

AI review should **add** to human review, not replace it. The most useful framing: the AI is a tireless junior reviewer who reads every line; the human is the senior reviewer who judges intent, design and risk.

## Configuring AI review so people read it

Noise kills AI review. If every PR gets twenty nitpicks, developers learn to ignore all of them, including the one real bug. Tune for signal:

- **Give it your standards.** Most tools read repository instructions (for example Copilot's instruction files, Bugbot rules, or AGENTS.md/CLAUDE.md for agents running in CI). Tell it what matters and what to ignore.
- **Ask for severity and confidence.** Only surface high-severity or high-confidence findings as blocking comments.
- **Focus by path.** Money, auth, data deletion and public APIs get deeper review.

```markdown
## Review guidelines (in AGENTS.md or the tool's review instructions)
- Prioritize: correctness, security, data loss, money/tax calculations, public API changes.
- Ignore: formatting (Prettier enforces), import order, naming nits unless misleading.
- For each finding give: severity (blocker/major/minor), confidence (high/medium/low),
  the exact line, and a concrete fix. Post at most 5 findings; summarize the rest.
- Flag any change to tests that weakens assertions.
```

## How humans should review agent-written code

Agent code looks fluent, which makes it easy to approve. Use a deliberate checklist:

1. **Intent first.** Read the issue and plan, then the PR description. Does the change do what was asked, and only that?
2. **Tests before code.** Read the tests: do they express the requirement? Are any weakened or skipped?
3. **Look for plausible fakes.** Hallucinated APIs, functions that exist in a different version, config keys that do nothing, invented environment variables.
4. **Check dependencies.** New packages? Are they real, maintained, and the intended ones (not typo-squats)?
5. **Security scan.** Input validation, authorization checks, secrets, logging of sensitive data.
6. **Duplication.** Agents often re-implement helpers that already exist. Ask "did we already have this?"
7. **Run it.** For anything user-facing, check out the branch and try it.

## Worked example: a review that caught the real bug

A Lahore fintech's AI reviewer flagged four minor issues on an agent's PR that added a refund endpoint. The human reviewer, following the checklist, read the tests first and noticed there was no test for a refund larger than the original payment. The implementation allowed it. Neither the agent nor the AI reviewer had the business rule, because it lived only in a finance team's spreadsheet. The fix: add the rule to the domain code, a test, and one line in AGENTS.md about refund limits.

## Hands-on: a self-review step before the PR

Ask the implementing agent to review its own diff with a fresh context (a subagent or new session) before opening the PR:

```text
You are reviewing a diff you did not write. Read `git diff main...HEAD`.
Report only: (1) bugs, (2) security issues, (3) requirement mismatches versus ISSUE.md,
(4) weakened or missing tests, (5) new dependencies. For each: file:line, severity, fix.
If you find nothing significant, say so. Do not comment on style.
```

A fresh context avoids the author's blind spots, and catching issues before the PR saves reviewer time.

## Pitfalls

- **Rubber-stamping fluent code.** Confidence of prose is not correctness of logic.
- **Letting AI approve PRs.** Keep a human approval requirement on protected branches.
- **Unbounded comment volume.** Tune or reviewers will mute the bot.

## How to measure success

Track the share of AI review comments that lead to a code change (a signal-to-noise proxy), defects found after merge, and median review time for agent PRs.

## Video lecture: Code review with AI, and reviewing AI code

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Code review with AI
2. Analogy: spell-checker and lawyer
3. Two directions
4. Strong vs weak
5. Tune for signal
6. Human checklist
7. Case: the missing rule
8. Self-review with fresh eyes
9. Rules and metrics
10. Example: the date parser
11. Scenario: tuning for signal (illustrative)
12. Deeper: two tuning lines that mattered
13. Watch me do it: reviewing a refund PR
14. Recap

## Lecture transcript

### Code review with AI

As agents write more of your code, review becomes the bottleneck and the last line of defense. In this lecture you will learn what AI reviewers are genuinely good at, how to tune them so developers actually read their comments, and a checklist for humans reviewing agent-written code, which is fluent, confident and occasionally very wrong.

### Analogy: spell-checker and lawyer

Here is an analogy for AI review. Think of proofreading a legal contract. A spell-checker catches every typo, every missing comma, every inconsistent defined term, instantly and without fatigue. But it cannot tell you that the contract gives away your intellectual property. For that you need a lawyer who understands what the deal is supposed to be. The AI reviewer is the world's best spell-checker for code, and more. The human reviewer is the lawyer. You need both, and you should never confuse their jobs.

### Two directions

AI review cuts both ways. First, AI reviewing code: tools like Copilot code review, Cursor's Bugbot, or Claude Code and Codex running in CI give every pull request a first pass. Second, humans reviewing code that agents wrote. Both matter, and they need different techniques. A useful mental model: the AI is a tireless junior reviewer who reads every line. The human is the senior reviewer who judges intent, design and risk.

### Strong vs weak

What is AI review strong at? Local bugs visible in the diff, consistency with nearby code, missing error handling and input validation, common security patterns like injection or leaked secrets, and summarizing big diffs. What is it weak at? Whether the change solves the right problem, architectural fit across services, business rules that are not written down, subtle concurrency issues, and the trade-offs your team has consciously accepted.

### Tune for signal

Noise is what kills AI review. If every pull request collects twenty nitpicks, developers learn to ignore the bot, including the one comment that mattered. So tune for signal. Give the reviewer your standards in its instruction file. Tell it what to prioritize, such as correctness, security, money and public APIs, and what to ignore, like formatting your linter already handles. Ask for severity and confidence on every finding, and cap the number of comments.

### Human checklist

Now the human side. Agent code reads fluently, which makes it dangerously easy to approve. Use a checklist. Intent first: does it do what the issue asked, and only that? Read the tests before the code. Look for plausible fakes, like APIs from a different library version or config keys that do nothing. Check any new dependency is real, maintained and not a typo-squat. Scan for security issues. Look for duplicated helpers. And run it.

### Case: the missing rule

Here is why the checklist matters. A fintech in Lahore had an agent add a refund endpoint. The AI reviewer flagged four minor issues. The human reviewer read the tests first and noticed something missing: no test for a refund larger than the original payment. The code allowed it. Neither the agent nor the AI reviewer could know the rule, because it lived only in the finance team's spreadsheet. The fix was a domain rule, a test, and one new line in the instruction file.

### Self-review with fresh eyes

A powerful habit is self-review with fresh eyes. Before opening a pull request, have a separate agent session, or a subagent, review the diff as if someone else wrote it. Ask only for bugs, security issues, mismatches with the issue, weakened tests and new dependencies, with file, line, severity and fix. A fresh context avoids the author's blind spots and saves your human reviewers time.

### Rules and metrics

Three rules keep the system honest. Never rubber-stamp fluent code. Never let an AI approve a pull request on a protected branch; keep a human approval requirement. And keep comment volume bounded. To track progress, measure how often AI review comments lead to a real code change, how many defects are found after merge, and the median review time for agent pull requests.

### Example: the date parser

A simple example. An agent's pull request adds a function that reads a date string with a parsing library. The AI reviewer comments: this call can throw on invalid input and the error is not handled. Useful, and correct. The human reviewer, reading the tests first, asks a different question: why are we parsing strings at all, when the upstream API already returns a timestamp? Removing the parsing removes the bug class entirely. The AI caught the local problem. The human caught the design problem. That division of labor is what good review looks like.

### Scenario: tuning for signal (illustrative)

Now a realistic scenario with illustrative numbers. A SaaS company in London turns on AI review for all repositories with default settings. In the first month it posts roughly twelve comments per pull request, and developers act on about one in ten. They tune it: priorities in the instruction file, a cap of five findings, and only high-severity comments posted inline. The next month, comments drop to around three per pull request, and developers act on about half of them. Fewer comments, much more value. Common mistake: judging an AI reviewer by how much it says, rather than how often it is right.

### Deeper: two tuning lines that mattered

One level deeper on tuning. The London team's biggest single improvement came from one instruction: ignore formatting and import order, because Prettier and the linter enforce them. That removed roughly half the comments overnight. The second came from requiring a concrete fix with every finding, which filtered out vague observations the reviewer could not justify.

### Watch me do it: reviewing a refund PR

Watch me do it. I'm reviewing an agent's pull request that adds a refund endpoint, using the fresh-eyes self-review first and then the human checklist. I open a new session and paste the self-review prompt: review the diff against ISSUE dot M D, report only bugs, security issues, requirement mismatches, weakened or missing tests and new dependencies, with file, line, severity and a fix. It returns three findings. A missing check that the order belongs to the caller: high severity. An unhandled error when the payment provider times out: medium. And a new dependency for date parsing that duplicates a helper we already have: low. The agent fixes all three before the pull request opens. Now the human pass. Intent: the issue asked for refunds up to the original amount; the pull request does that, and nothing else. Tests first: I read the test file and see happy path, unauthorized user and provider timeout. But there is no test for a refund larger than the payment. I check the code: the amount is validated against the order total, good, so I ask for the missing test anyway, because the rule should be pinned. Plausible fakes: the provider client method it calls exists in our installed version; I check the lock file to be sure. Dependencies: none new after the fix. Then I run it locally. Twelve minutes total, and I actually trust what I am approving.

### Recap

Recap. AI review is a tireless first pass, strongest on local issues. Tune it for signal with priorities, severity and caps. Humans review agent code with a checklist that starts with intent and tests. Your next step: add the review guidelines block from the lesson to your repository, and try the fresh-eyes self-review prompt on your next agent pull request.

## Key takeaways

- AI review is a tireless first pass, strongest on local, visible issues.
- Tune AI review for signal: priorities, ignores, severity, confidence and a comment cap.
- Review agent code with a checklist: intent, tests first, plausible fakes, dependencies, security, duplication, run it.
- Keep human approval on protected branches; measure how often AI comments lead to changes.

## Try it

Add review guidelines to your AGENTS.md and run the fresh-eyes self-review prompt on your next agent PR; note what it caught.

- [Previous: Debugging with agents: a scientific protocol](https://optimizeall.com/learn/agentic-coding-with-ai/debugging-with-agents)
- [Next: Refactoring and migrations at scale](https://optimizeall.com/learn/agentic-coding-with-ai/refactoring-and-migrations-at-scale)
- [All lessons of AI-Assisted Software Development: Coding Agents in Practice](https://optimizeall.com/learn/agentic-coding-with-ai)
