AI Security: Prompt Injection, Data Leakage and Red TeamingDefense in depth and red teaming · Lesson 15 of 17

Red-team tooling: garak, PyRIT, promptfoo and Inspect

Article · 14 min · 9 min lecture

Video lecture

Red-team tooling: garak, PyRIT, promptfoo and Inspect

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Red-team tooling

  • garak • PyRIT • promptfoo • Inspect
  • First scans
  • A practical pipeline

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The open-source red-team toolkit

Several mature open-source tools automate parts of AI red teaming. Each has a different sweet spot; most teams combine them. Versions and APIs move quickly, so pin versions and check each project's docs.

ToolFromSweet spot
garakNVIDIA (open source)Fast, broad vulnerability scanning of a model endpoint with many built-in probes (jailbreaks, encodings, prompt injection, leakage, toxicity) and detectors; CLI-first
PyRIT (Python Risk Identification Tool)Microsoft AI Red Team (open source)Programmable attack campaigns in Python: targets, converters (encodings, translations), scorers, multi-turn attack strategies, and memory of results
promptfoo red teamPromptfoo (open source; OpenAI announced an acquisition in March 2026 with a commitment to keep it open source)Application-level red teaming from config: describe the app's purpose, pick plugins (vulnerability types) and strategies (delivery techniques), run in CI, get a report mapped to frameworks such as OWASP
InspectUK AI Security InstituteRigorous evaluations including agentic tasks in sandboxes; good for building custom, reproducible attack evaluations

Other options include DeepTeam and commercial platforms. Choose based on whether you need breadth scanning, programmable campaigns, CI integration, or agentic sandboxes.

garak: a first scan

python -m pip install -U garak
export OPENAI_API_KEY=...   # or the relevant provider key; use a test key with a spending cap
python -m garak --target_type openai --target_name <model-id> --probes encoding
python -m garak --list_probes       # see available probe families

garak writes a report of which probes succeeded and which detectors fired. Use it to baseline a model or endpoint and to compare models; follow its docs to point it at a custom REST endpoint for your own application.

promptfoo: application red team in CI

# promptfooconfig.yaml (red team section) — verify plugin and strategy names in current docs
description: Souq & Style support agent red team
targets:
  - id: https
    label: support-agent-staging
    config:
      url: https://staging.souqstyle.example/api/chat
      method: POST
      headers:
        Content-Type: application/json
      body:
        message: "{{prompt}}"
      transformResponse: json.reply
redteam:
  purpose: >-
    Customer support agent for an online fashion retailer in PK, AE, SA and GB.
    It can look up orders only for verified customers and create refunds up to a limit.
    It must never reveal other customers' data or internal policies.
  numTests: 5
  plugins:
    - pii
    - bola               # broken object-level authorization (other users' objects)
    - bfla               # broken function-level authorization
    - rbac
    - excessive-agency
    - hijacking
  strategies:
    - jailbreak
    - prompt-injection
    - crescendo
    - base64
npx promptfoo@latest redteam run     # generate adversarial tests and run them against the target
npx promptfoo@latest redteam report  # open the findings report

For RAG apps, promptfoo documents an indirect-prompt-injection plugin that injects payloads into the variable holding retrieved context; configure it to match your prompt template.

PyRIT: programmable campaigns

PyRIT models red teaming as components you compose in Python:

  • Targets: your application endpoint or model.
  • Converters: transform attack prompts (Base64, character swaps, translation to other languages, audio/image conversion).
  • Scorers: decide whether an attempt succeeded (classifiers, LLM judges, substring checks).
  • Attack strategies / orchestration: single-turn batches, and multi-turn strategies where an adversarial model iterates toward an objective.
  • Memory: a database of all attempts and scores for analysis.

Its API has evolved across releases; start from the current documentation's quickstart and examples, and pin the version you use.

Building a practical pipeline

  1. Baseline scan (garak) on candidate models before selection or upgrade.
  2. Application red team in CI (promptfoo or custom Inspect/pytest suites) on every significant change, with confirmed findings as blocking tests.
  3. Campaigns (PyRIT or custom) for high-risk objectives, multilingual attacks and multi-turn strategies, run before major releases.
  4. Manual sessions focused on business logic and agent environments.
  5. Triage: automated findings include false positives; a human confirms, rates severity and writes the regression test.

Safety and hygiene for red-team tooling

  • Use staging targets and test API keys with spending caps; automated tools can generate thousands of calls.
  • Store attack corpora and results in restricted repositories; they contain harmful content.
  • Do not point tools at third-party systems without written authorization.
  • Log runs with tool version, target version and date for reproducibility.

Worked example

A UK retailer's team ran garak against two candidate models before an upgrade; one showed notably weaker resistance to encoding-based probes. They then ran promptfoo's red team in CI against the staging app, which found a broken object-level authorization issue (order lookup by guessed ID) that no model upgrade would have fixed. Manual testing added two business-logic findings. All confirmed findings became blocking tests.

Pitfalls

  • Trusting automated pass rates as "secure". Tools cover known patterns; novel attacks need humans.
  • Running scanners only against the base model.
  • No human triage, leading to false-positive fatigue or missed severity.
  • Unbounded spend from automated campaigns.

How to measure success

Scanner baselines for every model you deploy, an application red-team job in CI with blocking findings, scheduled campaigns for high-risk objectives, and a triage process that converts confirmed findings into regression tests.

Key takeaways

  • garak: fast CLI scanning of model endpoints with many probes and detectors.
  • PyRIT: programmable Python campaigns with targets, converters, scorers and multi-turn strategies.
  • promptfoo red team: config-driven application testing with plugins and strategies, CI-friendly reports.
  • Combine tools with manual testing and human triage; use staging, capped keys and restricted storage.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which tool is best described as config-driven application red teaming with plugins and strategies that runs in CI?
  2. A scanner shows a high pass rate. What should you conclude?

Put it into practice

Run a garak encoding scan against a test model with a capped key, and configure a promptfoo red-team job against a staging app with at least four plugins.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.