Skip to content

Building AI Products & Workflows · Data, privacy and security architecture · lesson 9 of 18 · 12 min

Data and privacy architecture for AI features

Map the data flow first

Every AI feature moves data: from users and systems, into prompts, to model providers, into logs and back out as outputs. Privacy and security problems usually come from flows nobody mapped. Start with a simple diagram:

User input --> App backend --> [Retrieval: internal docs, CRM] --> Prompt assembly
     --> Model provider API --> Output checks --> User
     \--> Logs / analytics / evaluation store (retention: ?)

For each arrow, ask: what data travels, how sensitive is it, where is it processed and stored, for how long, and who can access it?

Classify your data

Use your organisation's data classification (for example public, internal, confidential, restricted) and decide which classes may flow to which AI services. Personal data needs particular care, and special categories (health, biometric, financial details, children's data and similar) carry stricter rules in many jurisdictions.

Privacy principles applied to AI

  • Lawful basis and transparency: tell users how AI processes their data; ensure you have a lawful basis under applicable laws (for example UK GDPR and EU GDPR, and data protection laws in the UAE, Saudi Arabia, Pakistan and elsewhere as applicable).
  • Data minimisation: send only what the task needs. Often an ID and a few fields suffice; redact the rest.
  • Purpose limitation: don't reuse data collected for one purpose to train or tune for another without appropriate grounds.
  • Retention: set retention for prompts, outputs and logs; delete on schedule.
  • Rights: be able to find and delete a person's data, including in logs, evaluation sets and memory stores.
  • Cross-border transfers: know where your providers process data and whether transfer rules apply; some providers offer regional processing options.

This is not legal advice; involve your privacy or legal team early, especially for customer-facing features.

Provider terms to check

  • Is your data used to train the provider's models? (Many business and API offerings default to no, but verify.)
  • Retention periods for inputs and outputs, and options for shorter retention.
  • Data residency and sub-processors.
  • Security certifications and audit reports.
  • Contractual terms such as a data processing agreement.

Architecture patterns that reduce risk

  1. Redaction gateway: detect and mask personal data (names, emails, IDs) before prompts leave your environment, and re-insert where needed in outputs.
  2. Permission-aware retrieval: retrieval filters by the requesting user's access rights, so the AI never sees documents the user couldn't open.
  3. Tenant isolation: in multi-client products, strict separation of each client's data, indexes, memories and logs.
  4. Secrets outside prompts: API keys and credentials never appear in prompts or logs.
  5. Private or self-hosted models for the most sensitive workloads, where justified by risk and supported by operational capacity.
  6. Log hygiene: log what you need for debugging and evaluation, with masking and access controls.

Worked example: HR assistant

A company builds an assistant for employees' HR questions.

  • Data flow mapping shows employee questions may include health and family details.
  • Decisions: permission-aware retrieval over HR policies only (no individual records in v1); a redaction step for personal identifiers in logs; logs retained for a short defined period for quality review; a clear notice explaining how questions are processed; and escalation to a human HR partner for sensitive topics.
  • The provider's business terms confirm no training on inputs and offer regional processing.
  • Privacy team signs off on a data protection impact assessment before launch.

Common failure modes

  • Retrieval that ignores permissions, letting users "ask" their way into confidential documents.
  • Full prompts, including personal data, stored indefinitely in logs.
  • Evaluation datasets built from production data without review.
  • Assuming "enterprise" plans settle all privacy questions.

Hands-on: a minimal redaction gateway

A redaction step replaces personal identifiers with placeholders before text leaves your environment, keeps the mapping locally, and restores values in the output only where needed. Pattern-based redaction is a baseline, not a guarantee: combine it with named-entity detection for names and addresses, and test it on your real data (including Arabic and Urdu text).

import re, uuid

PATTERNS = {
    "EMAIL": re.compile(r"[\w.+-]+@[\w-]+\.[\w.-]+"),
    "PHONE": re.compile(r"(?<!\d)(?:\+?\d[\d\s-]{7,14}\d)(?!\d)"),
    "EMIRATES_ID": re.compile(r"\b784-?\d{4}-?\d{7}-?\d\b"),        # UAE ID format
    "CNIC": re.compile(r"\b\d{5}-\d{7}-\d\b"),                        # Pakistan CNIC format
    "CARD": re.compile(r"\b(?:\d[ -]?){13,19}\b"),
}

def redact(text: str):
    mapping = {}
    for label, pattern in PATTERNS.items():
        def _sub(m):
            token = f"[{label}_{uuid.uuid4().hex[:6]}]"
            mapping[token] = m.group(0)
            return token
        text = pattern.sub(_sub, text)
    return text, mapping

def restore(text: str, mapping: dict, allowed=("EMAIL",)):
    for token, value in mapping.items():
        if any(token.startswith(f"[{a}_") for a in allowed):
            text = text.replace(token, value)
    return text

safe, mapping = redact("Please call Ahmed on +971 50 123 4567 or email ahmed@example.com; ID 784-1985-1234567-1.")
print(safe)          # identifiers replaced before the model call
# ... send `safe` to the model, get `draft` back ...
# draft = restore(draft, mapping)    # restore only what the output genuinely needs

Pair it with log hygiene written as configuration, so it is reviewed like code:

logging:
  store_prompts: redacted_only        # never raw
  store_outputs: redacted_only
  retention_days: 30                  # agreed with privacy lead
  access: [ai-platform-oncall, quality-reviewers]
  evaluation_samples:
    source: production_logs
    requires: privacy_review
    max_rows_per_month: 500

Finally, add each AI feature to your record of processing (feature, purpose, data categories, providers and regions, retention, lawful basis, DPIA status). It is the document every customer security questionnaire and regulator will ask for.

Going further

Maintain a record of each AI feature's data flows, providers, data categories, retention and legal basis, updated with each significant change. It supports impact assessments, customer security questionnaires and regulatory inquiries, and it forces the right questions at design time.

Video lecture: Data and privacy architecture for AI features

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

  1. Data and privacy architecture
  2. Analogy: plumbing
  3. Map the flow
  4. Privacy principles for AI
  5. Provider terms checklist
  6. Risk-reducing patterns
  7. Simple example: tenant emails
  8. Worked example: HR assistant
  9. Business example (illustrative)
  10. Hands-on in the lesson
  11. Common mistakes
  12. How you'll know it's in shape
  13. Watch me do it: redaction + log policy
  14. Recap
  15. Try this now (30 minutes)

Lecture transcript

Data and privacy architecture

Every AI feature is a data pipeline in disguise. A customer's message goes into your backend, gets combined with documents from your CRM, travels to a model provider, comes back as an answer, and leaves copies in logs, analytics and evaluation sets along the way. Most AI privacy incidents come from flows nobody drew. In this lesson you'll map those flows, apply privacy principles to them, check provider terms, and build a redaction gateway.

Analogy: plumbing

Here's an analogy. Personal data in an AI feature is like water in a building's plumbing. You need to know where every pipe goes, which taps are for drinking water and which aren't, where the shut-off valves are, and where leaks could pool unseen, like under the floor. Logs and evaluation sets are those hidden pools. The data flow map is your plumbing diagram.

Map the flow

Start with a diagram. User input to your backend, retrieval from internal sources, prompt assembly, the model provider, output checks, back to the user, and the side branch into logs, analytics and evaluation stores. For every arrow, ask what data travels, how sensitive it is, where it's processed and stored, for how long, and who can access it. Then apply your data classification, public, internal, confidential, restricted, and decide which classes may flow to which AI services. Special categories like health, biometric and financial data carry stricter rules in many jurisdictions.

Privacy principles for AI

Now apply the privacy principles. Have a lawful basis and be transparent, under laws such as UK and EU GDPR and the data protection laws of the UAE, Saudi Arabia, Pakistan and elsewhere as applicable. Minimise: send only what the task needs, often an ID and a few fields. Limit purpose: don't reuse data collected for one purpose to train for another without grounds. Set retention for prompts, outputs and logs. Make sure you can find and delete a person's data, including in logs, evaluation sets and memory. And know where providers process data, since cross-border transfer rules may apply.

Provider terms checklist

Check provider terms carefully. Is your data used to train their models? Many business and API offerings default to no, but verify. What are the retention periods for inputs and outputs, and can you shorten them? Where is data processed, and who are the sub-processors? What certifications and audit reports exist? Is there a data processing agreement? Enterprise plan is not an answer to these questions, it's a reason to ask them.

Risk-reducing patterns

Six architecture patterns reduce risk. A redaction gateway masks personal data before prompts leave your environment and restores it only where needed. Permission-aware retrieval filters by the user's access rights, so the AI never sees what the user couldn't open. Tenant isolation separates each client's data, indexes, memories and logs. Secrets never appear in prompts or logs. Private or self-hosted models are used for the most sensitive workloads, where risk justifies the operational effort. And log hygiene: log what you need, masked, access-controlled and deleted on schedule.

Simple example: tenant emails

A simple example. A small estate agency adds an AI assistant that drafts replies to tenant emails. Tenants mention phone numbers, flat numbers and sometimes health issues affecting repairs. The team decides: redact phone numbers and ID numbers before the model call, keep tenant health details out of logs, keep logs for thirty days, and add a sentence to their privacy notice. Four decisions, made in an hour, before launch.

Worked example: HR assistant

Here's how it plays out. A company builds an assistant for employees' HR questions. Mapping shows questions may include health and family details. So version one retrieves only from HR policies, no individual records. Identifiers are redacted in logs, which are kept for a short, defined period for quality review. A clear notice explains how questions are processed, and sensitive topics escalate to a human HR partner. The provider's business terms confirm no training on inputs and offer regional processing, and the privacy team signs off a data protection impact assessment before launch.

Business example (illustrative)

Illustrative numbers for the HR assistant. In the first quarter it answered around four thousand questions. The redaction step masked identifiers in about one question in seven. Logs were deleted after thirty days as planned, and when one employee asked for their data to be removed, the team found and deleted it from logs and the evaluation sample in under an hour, because the data flow map listed every store.

Hands-on in the lesson

The hands-on section gives you a minimal redaction gateway in Python, with patterns for emails, phone numbers, UAE Emirates ID and Pakistani CNIC formats and card-like numbers, replacing each with a placeholder and keeping the mapping locally. It's a baseline, not a guarantee, so pair it with entity detection and test on real multilingual data. You'll also see log hygiene written as configuration, and the record of processing every AI feature should have.

Common mistakes

Common mistakes. Assuming the enterprise plan settles every privacy question. Logging full prompts with personal data forever, because debugging is easier. Building evaluation sets from production data without review. Retrieval that ignores document permissions. Forgetting to include logs and memories when someone asks for their data to be deleted. And involving the privacy team only a week before launch.

How you'll know it's in shape

How will you know your data architecture is in good shape? Every AI feature has an up-to-date data flow and record of processing. You can answer a customer security questionnaire in hours, not weeks. A deletion request covers logs, evaluation sets and memory, and you've tested it. And a spot check of logs finds no raw personal data where your policy says it shouldn't be.

Watch me do it: redaction + log policy

Watch me do it. I open the redaction gateway. First, the patterns: email, phone, the UAE Emirates ID format, the Pakistani CNIC format and card-like numbers. Next, redact loops through each pattern; for every match it creates a placeholder with the label and a short random ID, stores the original in a mapping, and substitutes it in the text. Restore puts back only the labels you allow, email by default. I run it on the sample sentence. The output replaces the phone number and ID with placeholders, and the mapping holds the originals locally, never sent to the provider. Then I open the logging config. Prompts and outputs are stored redacted only, retention is thirty days, access is limited to two groups, and evaluation samples need privacy review with a monthly cap. Finally, I add a row to the record of processing for this feature.

Recap

To recap: map every data flow, including logs and evaluation stores. Apply classification, lawful basis, minimisation, purpose limitation, retention and rights. Check provider terms properly. And use redaction, permission-aware retrieval, tenant isolation, secret hygiene and log discipline. This isn't legal advice, so involve your privacy team early. Your next step is to draw the data flow for one AI feature, annotate every arrow, and fix the riskiest one. Next: security threats and vendor risk.

Try this now (30 minutes)

Try this now. Pick one AI feature and draw its data flow on a single page, including logs, analytics and evaluation stores. Annotate every arrow with what data moves, its sensitivity, where it's processed, how long it's kept and who can access it. Circle the riskiest arrow and write one change that reduces the risk, like redaction or shorter retention. Share it with your privacy lead.

Key takeaways

  • Map every data flow: inputs, retrieval, provider, outputs, logs and evaluation stores.
  • Apply classification, lawful basis, minimisation, purpose limitation, retention and rights to AI data.
  • Check provider terms: training use, retention, residency, sub-processors and certifications.
  • Use redaction gateways, permission-aware retrieval, tenant isolation, secrets outside prompts and log hygiene.

Try it

Draw the data flow for one AI feature. For each arrow, note data sensitivity, location, retention and access, and identify one risk to fix.