AI Security: Prompt Injection, Data Leakage and Red TeamingOutput handling, data protection and cost abuse · Lesson 12 of 17

Secrets, PII, hidden context and unbounded consumption

Article · 14 min · 9 min lecture

Video lecture

Secrets, PII, hidden context and unbounded consumption

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Secrets, PII, hidden context and cost

  • Sensitive information disclosure
  • Hidden context exposure
  • Unbounded consumption

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Three ways data escapes without an "attack"

Some of the most common LLM security failures need no clever exploit: the system simply shows, stores or spends more than it should. This lesson covers Sensitive Information Disclosure (LLM02), Hidden Context Exposure (LLM08:2026, formerly System Prompt Leakage) and Unbounded Consumption (LLM06:2026).

Sensitive information disclosure

Sources of leaks:

  • Context over-sharing: retrieval or tools pull in more data than the question needs (the whole customer record when only the order status was asked).
  • Cross-user leakage: shared caches, conversation memory, or indexes without tenant isolation.
  • Training and fine-tuning data: models can memorize and regurgitate rare strings they were trained on; never fine-tune on secrets or unnecessary personal data.
  • Logs and traces: prompts and outputs captured in observability tools, analytics or error trackers.
  • Third-party processors: every model provider, judge model, and SaaS tool in the chain receives data.

Controls:

  1. Data minimization. Retrieve only fields needed; mask identifiers in context (show last four digits); summarize rather than paste.
  2. PII detection and redaction on inputs, context, outputs and telemetry (for example Microsoft Presidio plus custom recognizers for local formats such as CNIC, Emirates ID, Saudi national ID and IBANs).
  3. Tenant and user isolation in retrieval, caches and memory.
  4. Vendor terms. Retention, training use, region and zero-data-retention options for your plan; data processing agreements.
  5. Legal alignment. UK GDPR, EU GDPR, the UAE and Saudi PDPLs and Pakistan's emerging data protection regime set expectations on purpose limitation, minimization and cross-border transfer. Involve your privacy team.

Hidden context exposure

Everything your app puts in front of the model that users cannot see (system prompt, tool definitions, retrieved documents, memory, configuration) should be treated as potentially visible to users. Determined users can usually extract it, and the 2026 OWASP update widened this risk beyond the system prompt.

Rules:

  • No secrets in prompts or tool descriptions. API keys, connection strings, internal URLs with tokens, discount codes.
  • No security logic in prompts. "Only admins may see salaries" must be enforced by the tool, not the prompt.
  • Assume competitors can read your prompt. If that would hurt, move the value into code, data or fine-tuning, and accept the rest.
  • Detect leaks with canary strings in hidden context and monitoring for them in outputs.

Unbounded consumption

LLM calls cost money and capacity. Attackers (or bugs) can drive:

  • Denial of wallet: scripted requests to expensive models, very long inputs, requests for maximum-length outputs.
  • Agent loops: repeated tool calls, recursive delegation between agents.
  • Resource exhaustion: large file uploads, huge context assembly, expensive embeddings.
  • Model extraction: high-volume querying to copy behavior.

Controls:

ControlExample
Authentication and quotasPer-user and per-IP rate limits; daily token budgets per account tier
Input limitsMax characters, file size, pages, audio length
Output limitsmax_tokens per route
Agent limitsMax steps, max tool calls, max wall-clock time, loop detection
Cost alertsBudget alerts per API key and project; anomaly detection on spend
Model routingCheaper models for simple routes; expensive models behind auth
Abuse detectionFlag accounts with unusual volumes or patterns
# token budget guard (sketch): enforce before calling the model
from datetime import date
def check_budget(redis, user_id: str, requested_tokens: int, daily_limit: int = 200_000):
    key = f"tok:{user_id}:{date.today().isoformat()}"
    used = int(redis.get(key) or 0)
    if used + requested_tokens > daily_limit:
        raise PermissionError("daily AI budget reached")
    redis.incrby(key, requested_tokens)
    redis.expire(key, 60 * 60 * 26)

(Reconcile with actual usage from the provider's response afterwards; the illustrative limit should be set per tier.)

Worked example

A free AI writing tool launched by a small studio in Lahore was found by a scraper within days. Scripts sent thousands of long requests to the most expensive model through the unauthenticated endpoint, and the monthly bill arrived before anyone noticed. The fixes: authentication, per-account daily token budgets, input and output limits, a cheaper model for anonymous trials, provider-side budget alerts, and a status page message explaining the new limits.

Pitfalls

  • Secrets or discount codes in system prompts.
  • Logging full prompts with personal data to third-party tools without agreements.
  • Unauthenticated endpoints to expensive models.
  • No per-agent step limits.

How to measure success

No secrets in any hidden context (verified by scanning prompts and tool definitions in CI), PII redaction on telemetry, and hard budgets that stop cost spikes within minutes.

Key takeaways

  • Minimize data in context, mask identifiers, and redact PII in inputs, outputs and telemetry.
  • Treat all hidden context (prompts, tool definitions, retrieved docs, memory) as potentially visible; keep secrets and security logic out.
  • Use canary strings to detect hidden-context leaks.
  • Prevent denial of wallet with auth, quotas, token budgets, input/output/agent limits, alerts and model routing.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Where should a rule like 'only admins may see salaries' be enforced?
  2. An unauthenticated demo endpoint calls your most expensive model. What is the biggest risk?

Put it into practice

Scan all prompts and tool definitions for secrets, add a canary string, redact PII in telemetry, and set per-key budget alerts.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.