---
title: "Red-team tooling: garak, PyRIT, promptfoo and Inspect"
description: "The open-source red-team toolkit Several mature open-source tools automate parts of AI red teaming. Each has a different sweet spot; most teams combine…"
url: https://optimizeall.com/learn/ai-security-and-red-teaming/red-team-tools
updated: 2026-10-05
---

AI Security: Prompt Injection, Data Leakage and Red Teaming · Defense in depth and red teaming · lesson 15 of 17 · 14 min

# Red-team tooling: garak, PyRIT, promptfoo and Inspect

## The open-source red-team toolkit

Several mature open-source tools automate parts of AI red teaming. Each has a different sweet spot; most teams combine them. Versions and APIs move quickly, so pin versions and check each project's docs.

| Tool | From | Sweet spot |
|---|---|---|
| **garak** | NVIDIA (open source) | Fast, broad vulnerability *scanning* of a model endpoint with many built-in probes (jailbreaks, encodings, prompt injection, leakage, toxicity) and detectors; CLI-first |
| **PyRIT** (Python Risk Identification Tool) | Microsoft AI Red Team (open source) | *Programmable* attack campaigns in Python: targets, converters (encodings, translations), scorers, multi-turn attack strategies, and memory of results |
| **promptfoo red team** | Promptfoo (open source; OpenAI announced an acquisition in March 2026 with a commitment to keep it open source) | *Application-level* red teaming from config: describe the app's purpose, pick plugins (vulnerability types) and strategies (delivery techniques), run in CI, get a report mapped to frameworks such as OWASP |
| **Inspect** | UK AI Security Institute | Rigorous evaluations including agentic tasks in sandboxes; good for building custom, reproducible attack evaluations |

Other options include DeepTeam and commercial platforms. Choose based on whether you need breadth scanning, programmable campaigns, CI integration, or agentic sandboxes.

## garak: a first scan

```bash
python -m pip install -U garak
export OPENAI_API_KEY=...   # or the relevant provider key; use a test key with a spending cap
python -m garak --target_type openai --target_name <model-id> --probes encoding
python -m garak --list_probes       # see available probe families
```

garak writes a report of which probes succeeded and which detectors fired. Use it to baseline a model or endpoint and to compare models; follow its docs to point it at a custom REST endpoint for your own application.

## promptfoo: application red team in CI

```yaml
# promptfooconfig.yaml (red team section) — verify plugin and strategy names in current docs
description: Souq & Style support agent red team
targets:
  - id: https
    label: support-agent-staging
    config:
      url: https://staging.souqstyle.example/api/chat
      method: POST
      headers:
        Content-Type: application/json
      body:
        message: "{{prompt}}"
      transformResponse: json.reply
redteam:
  purpose: >-
    Customer support agent for an online fashion retailer in PK, AE, SA and GB.
    It can look up orders only for verified customers and create refunds up to a limit.
    It must never reveal other customers' data or internal policies.
  numTests: 5
  plugins:
    - pii
    - bola               # broken object-level authorization (other users' objects)
    - bfla               # broken function-level authorization
    - rbac
    - excessive-agency
    - hijacking
  strategies:
    - jailbreak
    - prompt-injection
    - crescendo
    - base64
```

```bash
npx promptfoo@latest redteam run     # generate adversarial tests and run them against the target
npx promptfoo@latest redteam report  # open the findings report
```

For RAG apps, promptfoo documents an `indirect-prompt-injection` plugin that injects payloads into the variable holding retrieved context; configure it to match your prompt template.

## PyRIT: programmable campaigns

PyRIT models red teaming as components you compose in Python:

- **Targets:** your application endpoint or model.
- **Converters:** transform attack prompts (Base64, character swaps, translation to other languages, audio/image conversion).
- **Scorers:** decide whether an attempt succeeded (classifiers, LLM judges, substring checks).
- **Attack strategies / orchestration:** single-turn batches, and multi-turn strategies where an adversarial model iterates toward an objective.
- **Memory:** a database of all attempts and scores for analysis.

Its API has evolved across releases; start from the current documentation's quickstart and examples, and pin the version you use.

## Building a practical pipeline

1. **Baseline scan** (garak) on candidate models before selection or upgrade.
2. **Application red team in CI** (promptfoo or custom Inspect/pytest suites) on every significant change, with confirmed findings as blocking tests.
3. **Campaigns** (PyRIT or custom) for high-risk objectives, multilingual attacks and multi-turn strategies, run before major releases.
4. **Manual sessions** focused on business logic and agent environments.
5. **Triage:** automated findings include false positives; a human confirms, rates severity and writes the regression test.

## Safety and hygiene for red-team tooling

- Use **staging targets** and **test API keys with spending caps**; automated tools can generate thousands of calls.
- Store attack corpora and results in **restricted repositories**; they contain harmful content.
- Do not point tools at **third-party systems** without written authorization.
- Log runs with tool version, target version and date for reproducibility.

## Worked example

A UK retailer's team ran garak against two candidate models before an upgrade; one showed notably weaker resistance to encoding-based probes. They then ran promptfoo's red team in CI against the staging app, which found a broken object-level authorization issue (order lookup by guessed ID) that no model upgrade would have fixed. Manual testing added two business-logic findings. All confirmed findings became blocking tests.

## Pitfalls

- **Trusting automated pass rates as "secure".** Tools cover known patterns; novel attacks need humans.
- **Running scanners only against the base model.**
- **No human triage**, leading to false-positive fatigue or missed severity.
- **Unbounded spend** from automated campaigns.

## How to measure success

Scanner baselines for every model you deploy, an application red-team job in CI with blocking findings, scheduled campaigns for high-risk objectives, and a triage process that converts confirmed findings into regression tests.

## Video lecture: Red-team tooling: garak, PyRIT, promptfoo and Inspect

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Red-team tooling
2. Analogy: sensors and a consultant
3. garak (NVIDIA)
4. PyRIT (Microsoft)
5. promptfoo red team
6. Inspect (UK AISI)
7. Example: before an upgrade
8. The pipeline
9. Hygiene + a warning
10. promptfoo config, explained
11. Triage in practice
12. Deeper: controlling red-team spend
13. Watch me do it: garak + promptfoo
14. Recap

## Lecture transcript

### Red-team tooling

Manual red teaming finds the clever attacks. But you also need breadth: thousands of known attack variants, re-run every time your model, prompt or tools change. That is where open-source tools come in. In this lecture you will meet four of the most used, garak, PyRIT, promptfoo's red team features and Inspect, learn what each is best at, run a first scan, and assemble them into a practical pipeline.

### Analogy: sensors and a consultant

Why automate at all? Think of home security again. A security consultant walking your house once a year is valuable; they notice the loose window latch nobody else would. But you also want sensors on every door and window, checking every minute. Automated red-team tools are the sensors: tireless, broad, repeatable. Humans are the consultant. You need both, and the tools only earn their keep if someone reviews what they find.

### garak (NVIDIA)

Tool one: garak, from NVIDIA. It is a command line vulnerability scanner for model endpoints, with many built-in probes, covering jailbreaks, encodings, prompt injection, data leakage and toxicity, plus detectors that judge whether each probe worked. It is excellent for fast, broad baselines, for example comparing two candidate models before an upgrade. The lesson shows a first scan: install it, set a test key with a spending cap, and run the encoding probes against a model.

### PyRIT (Microsoft)

Tool two: PyRIT, the Python Risk Identification Tool from Microsoft's AI Red Team. Where garak scans, PyRIT orchestrates. You compose targets, which are your app or model; converters, which transform attack prompts through encodings, translations, or audio and image formats; scorers, which decide if an attempt succeeded; and attack strategies, including multi-turn ones where an adversarial model iterates toward an objective. Every attempt is stored in memory for analysis. Its API has evolved, so start from the current quickstart and pin your version.

### promptfoo red team

Tool three: promptfoo's red team features. You describe your application's purpose in plain language, choose plugins, which are vulnerability types like personal data leakage, broken object-level authorization, excessive agency and hijacking, and choose strategies, which are delivery techniques like jailbreak templates, prompt injection, multi-turn crescendo and Base64 encoding. It generates tailored attacks, runs them against your real endpoint, and produces a report mapped to frameworks like OWASP. And it runs nicely in CI. OpenAI announced its acquisition of Promptfoo in March twenty twenty-six and committed to keeping it open source.

### Inspect (UK AISI)

Tool four: Inspect, from the UK AI Security Institute. It is a rigorous evaluation framework with strong support for agent tasks inside Docker sandboxes. It is ideal when you want custom, reproducible attack evaluations, for example planting malicious content in a sandboxed file system and checking whether an agent exfiltrates anything. Other options exist too, including DeepTeam and commercial platforms. Choose based on what you need: breadth scanning, programmable campaigns, CI integration or agent sandboxes.

### Example: before an upgrade

A simple example of combining them. Before a model upgrade, you run garak against both the current and candidate models. The candidate turns out weaker against encoding-based probes. That is useful input, but not the whole story. Next you run promptfoo's red team against your staging application. It finds that a guessed order number returns someone else's order, a broken object-level authorization bug in your code. No model upgrade would ever have fixed that. The scanner judged the model. The application red team judged your system.

### The pipeline

Now the pipeline, as a realistic setup a retailer might run. Baseline scans with garak on every candidate model. An application red-team job in CI on every significant change, with confirmed findings as blocking tests. Scheduled PyRIT or custom campaigns before major releases for high-risk objectives, including multilingual and multi-turn attacks. Manual sessions focused on business logic and agent environments. And human triage, because automated findings include false positives, and someone must confirm, rate severity and write the regression test.

### Hygiene + a warning

Hygiene matters, because these tools are powerful. Point them at staging, not production, and use test API keys with spending caps, since automated campaigns can generate thousands of calls. Store attack corpora and results in restricted repositories, because they contain harmful content. Never aim tools at third-party systems without written authorization. And log every run with tool version, target version and date. Common mistake: treating an automated pass rate as proof of security. Tools cover known patterns. Humans find the new ones.

### promptfoo config, explained

Let's make the promptfoo example concrete. In the lesson's config, the purpose says the agent supports a fashion retailer in four countries, can look up orders only for verified customers, and can create refunds up to a limit. That one paragraph matters, because the tool uses it to generate attacks tailored to your business, like pretending to be a verified customer or asking for a refund split into smaller pieces. The plugins choose what to test for, such as personal data leakage and broken authorization. The strategies choose how to deliver each attack. Then you run it, open the report, and triage.

### Triage in practice

And here is how triage works in practice. The report shows, say, forty failed tests. A human reviews them. Some are false positives: the agent refused politely, but the grader misread it. Some are duplicates of one root cause, like five variations that all exploit the same missing ownership check. And one or two are genuinely new. The output of triage is short: a handful of confirmed findings, each with a severity, an owner and a regression test. Without triage, teams either drown in noise or ignore the report entirely.

### Deeper: controlling red-team spend

One level deeper on cost control. Before the promptfoo run, the team set the number of tests per plugin to five and capped the staging key at a small daily budget. The first run generated a few hundred attacks, which was plenty for a baseline. They increased coverage only for plugins that showed failures.

### Watch me do it: garak + promptfoo

Watch me do it: a garak scan, then a promptfoo red-team run. First, garak. I install it, set a test API key with a spending cap, and run python dash m garak with target type openai, our candidate model's name and the encoding probes. It sends encoded payloads, Base64, hex and others, and its detectors decide whether the model decoded and followed them. The report lists each probe with a pass rate. I run the same command against our current model and put the two reports side by side: the candidate resists slightly less on two encoding probes. That goes into the model-selection notes. Second, promptfoo against the staging app. The target block points at our staging chat endpoint, sends the prompt in the message field and reads the reply field. The purpose paragraph describes the retailer, verified-only order lookups and capped refunds. Plugins: personal data, broken object-level authorization, broken function-level authorization, role-based access control, excessive agency and hijacking. Strategies: jailbreak, prompt injection, crescendo and Base64. I run redteam run, then redteam report. The report shows most categories passing, but several failures under object-level authorization. Triage: four are the same root cause, order lookup by guessed ID. One confirmed finding, one owner, one regression test.

### Recap

Recap. Use garak for fast breadth baselines, PyRIT for programmable, multi-turn campaigns, promptfoo for application red teaming in CI, and Inspect for reproducible agent evaluations in sandboxes. Combine them with manual testing and human triage, and turn every confirmed finding into a regression test. Try this now: run a garak encoding scan against a test model with a capped key, then configure a promptfoo red-team job against your staging app with at least four plugins.

## Key takeaways

- garak: fast CLI scanning of model endpoints with many probes and detectors.
- PyRIT: programmable Python campaigns with targets, converters, scorers and multi-turn strategies.
- promptfoo red team: config-driven application testing with plugins and strategies, CI-friendly reports.
- Combine tools with manual testing and human triage; use staging, capped keys and restricted storage.

## Try it

Run a garak encoding scan against a test model with a capped key, and configure a promptfoo red-team job against a staging app with at least four plugins.

- [Previous: Planning and running AI red-team engagements](https://optimizeall.com/learn/ai-security-and-red-teaming/red-teaming-methods)
- [Next: Incident response for AI systems](https://optimizeall.com/learn/ai-security-and-red-teaming/ai-incident-response)
- [All lessons of AI Security: Prompt Injection, Data Leakage and Red Teaming](https://optimizeall.com/learn/ai-security-and-red-teaming)
