AI-Assisted Software Development: Coding Agents in PracticeCode review, migrations and CI integration · Lesson 12 of 17

Integrating agents into CI with GitHub Actions

Article · 14 min · 9 min lecture

Video lecture

Integrating agents into CI with GitHub Actions

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Agents in CI

  • What the vendors offer
  • CI security rules
  • Review, autofix and triage patterns

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Agents inside your pipeline

Running agents in CI turns them from personal productivity tools into team infrastructure: every PR gets a first review, failing builds get a proposed fix, issues get triaged. All the major vendors support it:

  • Claude Code GitHub Action (anthropics/claude-code-action) responds to @claude mentions or runs a fixed prompt in any workflow; it can also run headless (claude -p) in any CI.
  • Codex GitHub Action (openai/codex-action) runs codex exec with configurable sandbox and safety strategy.
  • GitHub Copilot cloud agent works from assigned issues in its own GitHub Actions-powered environment, customizable with a copilot-setup-steps.yml workflow; Copilot code review can be requested automatically on PRs.
  • Cursor's Bugbot and other review apps install as GitHub apps.

Security first: CI is a high-value target

A CI runner often holds secrets and write tokens. Adding an agent that reads untrusted text (PR titles, issue bodies, code from forks) is a textbook prompt-injection risk. Rules:

  1. Least-privilege permissions: on every workflow. Start from contents: read and add only what the job needs.
  2. Do not run agents with secrets on untrusted fork PRs. Be very careful with pull_request_target, which runs with the base repository's secrets; never check out and execute fork code in that context.
  3. Restrict who can trigger mention-based agents (for example, only members with write access), and check the action's own settings for this.
  4. Limit tools and turns. Allow only the commands the job needs; set max turns and timeouts.
  5. Agents propose, humans merge. Output is a comment or a PR, never a direct push to protected branches.
  6. Pin actions to a full commit SHA or a trusted major version, and review updates.
  7. Keep API keys in repository or organization secrets, never in workflow files.

Hands-on: AI review on every PR (Claude Code Action)

# .github/workflows/ai-review.yml
name: ai-review
on:
  pull_request:
    types: [opened, synchronize]
permissions:
  contents: read
  pull-requests: write
jobs:
  review:
    if: github.event.pull_request.head.repo.full_name == github.repository  # skip forks
    runs-on: ubuntu-latest
    timeout-minutes: 15
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - uses: anthropics/claude-code-action@v1
        with:
          anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
          prompt: |
            Review this pull request using the review guidelines in AGENTS.md.
            Report at most 5 findings with severity, confidence, file:line and a fix.
            Post them as a single PR comment. Do not modify files.
          claude_args: |
            --max-turns 8

Check the action's README for current inputs and for how to restrict allowed tools.

Hands-on: propose a fix when CI fails (Codex Action)

# .github/workflows/ci-autofix.yml (runs after your main CI on the same repo only)
name: ci-autofix
on:
  workflow_run:
    workflows: ["ci"]
    types: [completed]
permissions:
  contents: write
  pull-requests: write
jobs:
  autofix:
    if: >-
      github.event.workflow_run.conclusion == 'failure' &&
      github.event.workflow_run.head_repository.full_name == github.repository
    runs-on: ubuntu-latest
    timeout-minutes: 20
    steps:
      - uses: actions/checkout@v4
        with:
          ref: ${{ github.event.workflow_run.head_sha }}
      - uses: openai/codex-action@v1
        with:
          openai-api-key: ${{ secrets.OPENAI_API_KEY }}
          prompt: "Identify the minimal change to make the failing tests pass. Do not edit tests."
          sandbox: workspace-write
          safety-strategy: drop-sudo
      # then: open a PR with the changes (e.g. peter-evans/create-pull-request), never push to main

Input names and defaults evolve; verify against the action's current documentation before use.

Copilot cloud agent environment

For Copilot's cloud agent, a .github/workflows/copilot-setup-steps.yml file with a single copilot-setup-steps job preinstalls your dependencies so the agent can build and test. Keep its built-in firewall enabled unless you have a strong reason and compensating controls; disabling it (for example for some self-hosted runner setups) moves the network trust boundary entirely onto you.

Worked example: triage bot for a creator-tools startup

A small team in London building tools for YouTube creators gets dozens of issues a week. A scheduled workflow runs an agent with read-only permissions plus issues: write to label new issues, detect duplicates, and ask for missing reproduction steps. It never closes issues. Maintainers review the labels weekly. Because issue text is untrusted, the agent has no shell access in that job, and its only write capability is adding labels and a comment.

Costs and observability

  • Set max turns and timeouts per job; agents stuck in loops burn money.
  • Log token usage per workflow run and set organization-level budget alerts.
  • Keep transcripts or logs for audit, with secrets masked.

Pitfalls

  • permissions: write-all "to make it work".
  • Agents on fork PRs with secrets.
  • Auto-merge of agent PRs without human approval.
  • Unpinned third-party actions.

How to measure success

Track proposals accepted versus rejected per job type, cost per accepted change, and any security findings from periodic reviews of workflow permissions.

Key takeaways

  • Claude Code, Codex, Copilot and review apps all integrate with GitHub; verify current inputs in each action's docs.
  • CI agents read untrusted text while holding credentials: apply least-privilege permissions and skip forks.
  • Agents propose (comments or PRs); humans merge. Never auto-merge to protected branches.
  • Bound cost with max turns, timeouts and budget alerts; pin actions and keep keys in secrets.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which workflow setting is safest for an AI review job?
  2. An agent triage bot reads new issue bodies. What capabilities should it have?

Put it into practice

Add the AI review workflow to a test repository, confirm it skips fork PRs, and document its permissions and cost limits.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.