---
title: "Codex: OpenAI’s coding agent for builders and teams"
description: "What Codex is Codex is OpenAI's coding agent. It can write features, fix bugs, refactor, write tests, answer questions about a codebase and propose pull…"
url: https://optimizeall.com/learn/mastering-chatgpt/codex-for-builders
updated: 2026-10-05
---

Mastering ChatGPT (OpenAI) · Apps, plugins and Codex · lesson 15 of 19 · 16 min

# Codex: OpenAI’s coding agent for builders and teams

## What Codex is

**Codex** is OpenAI's coding agent. It can write features, fix bugs, refactor, write tests, answer questions about a codebase and propose pull requests. It runs in several places, all tied to your ChatGPT plan (Plus, Pro, Business, Enterprise and Edu, with usage limits that vary):

- **Codex Web (cloud):** at chatgpt.com/codex, tasks run in cloud environments connected to your GitHub repositories, several in parallel.
- **Codex CLI:** a local agent in your terminal that reads, edits and runs code on your machine.
- **IDE extension:** in VS Code and compatible editors such as Cursor and Windsurf.
- **Codex app:** a desktop command centre for running and supervising many agent tasks, with built-in worktrees and cloud environments.
- **Code review:** Codex can review pull requests on GitHub.

OpenAI's newest coding models power Codex and are updated frequently.

## Getting started with the CLI

Install using the official instructions (standalone installer, npm or Homebrew), then run it in a project:

```bash
# one option from the official README
npm install -g @openai/codex

cd my-store-scripts
codex
```

Sign in with your ChatGPT account (or an API key for usage-based billing). Codex asks for approval according to the approval mode you choose; start conservative so it asks before running commands or editing outside the workspace.

## Guide it with AGENTS.md

Codex reads an **AGENTS.md** file in your repository for project guidance, similar to a README for agents:

```markdown
# AGENTS.md
## Project
Node.js scripts that sync Shopify orders to our Google Sheet and Slack.

## Commands
- Install: npm ci
- Test: npm test (must pass before any change is considered done)
- Lint: npm run lint

## Rules
- Never commit secrets; read tokens from environment variables.
- Don't modify /migrations or production config.
- Prefer small, focused changes with tests.
- Currency amounts are integers in minor units (fils/halalas/paisa).
```

Keep it short, specific and up to date.

## A delegation pattern that works

1. **Plan first:** "Read the repo and propose a plan to add retry logic to the Slack notifier. Don't change code yet."
2. **Small tasks:** one feature or fix per task, with a clear definition of done ("tests pass; new test covers a 429 response").
3. **Parallelise carefully:** in Codex Web or the app, run independent tasks in parallel on separate branches or worktrees.
4. **Review like any pull request:** read the diff, run the tests, check for secrets and unexpected dependencies.
5. **Use Codex code review** as an extra reviewer, not a replacement for a human.

## Security and governance

- Keep secrets out of prompts, AGENTS.md and code; use environment variables or a secrets manager.
- Limit what cloud environments can access (network access and credentials) to what the task needs.
- Decide who can merge AI-authored changes and require tests and review.
- On business plans, admins can manage Codex access and settings.

## Worked example: a Shopify agency's backlog

A three-person agency in Karachi maintains custom scripts for a dozen Shopify stores in the Gulf. They add AGENTS.md files, then use Codex Web to run five small backlog tasks in parallel overnight (currency formatting for SAR, a flaky test, a CSV export column). In the morning they review five pull requests: three merge after minor edits, one needs rework, one is rejected because it added an unnecessary dependency. Backlog time drops substantially, and every change still passes human review and tests.

## Codex vs other coding agents

Claude Code, GitHub Copilot's agent features, Google's tools and others offer similar agentic coding. The durable skills are the same: guidance files, plan-first delegation, small tasks, tests, and review. Choose based on your repository host, security requirements and which models perform best on your codebase in a fair trial.

## Hands-on

Install Codex (CLI or IDE extension) on a non-critical repository, add a short AGENTS.md, and complete one small task plan-first with a test. Review the diff as you would a colleague's.

## Using Codex from the API

Teams that build their own tooling can also use OpenAI's coding models through the API (Module 7), and the **Codex SDK** and **Agents SDK** let developers embed agentic coding or custom agents in their own workflows, for example a CI job that asks an agent to propose a fix when a test fails. Keep the same rules: least-privilege credentials, sandboxed execution, and human review before merge.

## Pitfalls

- Large, vague tasks ("improve the codebase").
- Merging without running tests or reading the diff.
- Giving cloud environments production credentials.
- Stale AGENTS.md files.

## How to measure success

Cycle time for small tasks falls, defect rates do not rise, and every AI-authored change is tested and reviewed before merging.

## Video lecture: Codex: OpenAI’s coding agent for builders and teams

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Codex
2. Why coding agents change the job
3. Where Codex runs
4. Getting started
5. AGENTS.md
6. Delegation pattern
7. Security and governance
8. Simple example
9. Worked example: a Karachi Shopify agency
10. Codex among coding agents
11. Pitfalls
12. Try this now
13. Watch me do it, part 1
14. Watch me do it, part 2
15. Recap and next step

## Lecture transcript

### Codex

Imagine giving five small coding tasks to a tireless teammate at six in the evening and reviewing five pull requests with your morning coffee. That is what Codex offers. In this lecture you will learn where Codex runs, how to guide it with an AGENTS dot M D file, a delegation pattern that produces mergeable work, and the governance that keeps it safe.

### Why coding agents change the job

Why do coding agents change the job? Because the work shifts from typing code to guiding and reviewing it. That is a big productivity gain, but only if the guidance is clear and the review is real. Think of it like managing a very fast junior developer. They can produce a lot in a short time, and your job becomes setting clear tasks, providing the house rules and checking the work before it ships.

### Where Codex runs

Codex is OpenAI's coding agent, and it runs in several places tied to your ChatGPT plan. Codex Web runs tasks in cloud environments connected to your GitHub repositories, several in parallel. The Codex CLI is a local agent in your terminal. There is an extension for VS Code and compatible editors. The Codex app is a desktop command centre for supervising many agents, with built in worktrees. And Codex can review pull requests on GitHub.

### Getting started

To start with the CLI, install it using the official instructions, which offer a standalone installer, npm or Homebrew. Change into your project folder and run codex. Sign in with your ChatGPT account, or use an API key if you want usage based billing. Choose a cautious approval mode at first, so Codex asks before running commands or editing outside your workspace.

### AGENTS.md

Codex reads a file called AGENTS dot M D, a readme for agents. Include a one line project summary, the commands to install, test and lint, and rules. Never commit secrets. Do not touch migrations or production config. Prefer small changes with tests. And local details, like storing currency amounts as integers in fils, halalas or paisa. Keep it short, specific and current.

### Delegation pattern

Here is the delegation pattern. Plan first, read the repo and propose a plan to add retry logic, do not change code yet. Then small tasks, one feature or fix each, with a definition of done such as tests pass and a new test covers a rate limit response. Run independent tasks in parallel on separate branches or worktrees. Review every change like a colleague's pull request. And use Codex code review as an extra reviewer, never as a replacement for a human.

### Security and governance

Governance matters. Keep secrets out of prompts, AGENTS dot M D and code, using environment variables or a secrets manager. Limit what cloud environments can reach, both network access and credentials, to what the task needs. Decide who may merge AI authored changes, and require tests and review. On business plans, admins can manage Codex access and settings.

### Simple example

A simple example. In a small project, ask Codex, add a unit test for the currency formatter that covers Saudi riyals and a zero amount. It reads the code, writes the test, runs the test suite and shows you the diff. The change is small enough to read in two minutes. You check that the test actually tests what you asked, run it yourself, and merge. That is the ideal size for a first task.

### Worked example: a Karachi Shopify agency

A three person agency in Karachi maintains custom scripts for a dozen Shopify stores in the Gulf. They add AGENTS dot M D files, then use Codex Web to run five small backlog tasks overnight, currency formatting for Saudi riyals, a flaky test, a new CSV export column. In the morning they review five pull requests. Three merge after minor edits, one needs rework, and one is rejected because it added an unnecessary dependency. Backlog time drops, and every change still passes human review and tests. They also learned what not to delegate. A task to redesign the whole order sync architecture produced a large, confident diff they could not review properly, so they rejected it and broke the work into six small tasks instead. Each small task was reviewed in minutes. Big changes still happen, just one reviewable step at a time.

### Codex among coding agents

Codex is not the only coding agent. Claude Code, GitHub Copilot's agent features and others offer similar capabilities. The durable skills are identical, a guidance file, plan first delegation, small tasks, tests and review. Choose based on where your code lives, your security requirements, and a fair trial on your own codebase.

### Pitfalls

Four pitfalls. Vague, huge tasks like improve the codebase. Merging without running tests or reading the diff. Giving cloud environments production credentials. And a stale AGENTS dot M D. Measure success by cycle time for small tasks falling while defect rates stay flat or improve.

### Try this now

Try this now. Install Codex, the CLI or the editor extension, and open a repository that is not critical, perhaps an internal script or a personal project. Write a ten line AGENTS dot M D with the test command and three rules. Ask Codex for a plan for one small improvement, approve it, and let it make the change with a test. Review the diff as you would a colleague's, request at least one change, and only then merge. Note how long the whole loop took.

### Watch me do it, part 1

Let me walk through the Karachi agency's first task. I install the Codex CLI using the official instructions and run codex inside the store scripts repository. First I write AGENTS dot M D. Project, Node scripts syncing Shopify orders to Google Sheets and Slack. Commands, npm ci, npm test, npm run lint. Rules, never commit secrets, do not touch migrations, small focused changes, and currency amounts are integers in minor units. Then I ask, propose a plan to add a unit test for the currency formatter covering Saudi riyals and a zero amount, do not change code yet. It proposes three steps, and I approve.

### Watch me do it, part 2

Codex writes the test and runs the suite. In the diff, I notice it also changed the formatter itself to make one case pass. That is not what I asked for, and it might hide a real bug. So I reply, do not change the formatter, only add tests, and report any failing case. The revised diff adds tests only, and one test fails, revealing that zero riyals prints as an empty string. That is a real bug, now documented. I open a separate small task to fix it, review that diff too, run the tests, and merge both. That is the rhythm, small tasks, real review.

### Recap and next step

Recap. Codex runs in the cloud, your terminal, your editor, a desktop app and GitHub reviews. Guide it with AGENTS dot M D, delegate plan first in small tasks with clear definitions of done, keep secrets and credentials locked down, and review everything. Your next step: install Codex on a non critical repository, write a short AGENTS dot M D, and complete one small task with a test.

## Key takeaways

- Codex is OpenAI’s coding agent across Codex Web (cloud), CLI, IDE extension, the Codex app and GitHub code review, tied to ChatGPT plans.
- An AGENTS.md file gives Codex project commands, rules and local conventions; keep it short and current.
- Delegate plan-first in small tasks with a definition of done; run independent tasks in parallel and review every diff.
- Keep secrets out of prompts and files, limit cloud environment access, and require tests and human review before merging.

## Try it

Install Codex (CLI or IDE extension) on a non-critical repository, write a 10-line AGENTS.md, and complete one small task plan-first with a new test. Review the diff and note one change you requested.

- [Previous: Apps, plugins and connected data in ChatGPT](https://optimizeall.com/learn/mastering-chatgpt/apps-and-plugins-in-chatgpt)
- [Next: The OpenAI API: Responses API fundamentals with working code](https://optimizeall.com/learn/mastering-chatgpt/openai-api-fundamentals)
- [All lessons of Mastering ChatGPT (OpenAI)](https://optimizeall.com/learn/mastering-chatgpt)
