Responsible AI, Disclosure & ComplianceData protection and your AI usage policy · Lesson 10 of 11

Data protection principles for AI use

Article · 16 min · 8 min lecture

Video lecture

Data protection principles for AI use

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Data protection for AI use

  • When it applies
  • Seven principles
  • AI-specific traps
  • Redact before you prompt
  • One-page DPIA

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why data protection applies to AI

Whenever AI processes information about identifiable people (customers, followers, leads, employees, creators), data protection law applies. That includes pasting a customer email into a chatbot, running an AI lead-scoring tool, transcribing sales calls, cloning a voice and building a lookalike audience.

The core principles

Most modern data protection laws, including the EU GDPR, UK GDPR, the UAE federal Personal Data Protection Law, Saudi Arabia's PDPL, laws in DIFC and ADGM, US state privacy laws, and data protection frameworks being developed in Pakistan and elsewhere, share principles like these:

  1. Lawfulness, fairness and transparency: have a valid legal basis (such as consent, contract or legitimate interests where recognized) and tell people how you use their data, including AI processing.
  2. Purpose limitation: use data only for the purposes you told people about. Customer support emails collected to answer queries should not quietly become training data for a marketing model.
  3. Data minimization: use only what you need. Remove names and contact details before AI analysis where possible.
  4. Accuracy: keep data correct, and be aware that AI-generated inferences about people can be wrong.
  5. Storage limitation: don't keep data, including chat logs and transcripts, longer than necessary.
  6. Security: protect data with access controls, business-grade tools and secure accounts.
  7. Accountability: be able to show what you do: records, policies and assessments.

Rights people have

Commonly: access to their data, correction, deletion, objection to certain processing (including direct marketing), and in some laws, safeguards around solely automated decisions with significant effects. If an AI system alone decides something significant about a person, such as rejecting an application, you may need human involvement, an explanation and a way to contest.

Key AI-specific issues

Sharing with AI providers. When you send personal data to an AI provider, they typically act as your processor or service provider. Under GDPR-style laws you generally need a data processing agreement covering security, confidentiality, sub-processors and use restrictions. Business plans from major providers usually offer these; consumer plans often do not.

Training on your data. Check whether the provider uses your inputs to train models, and opt out or use business tiers for personal and client data.

International transfers. Data sent to providers in other countries may trigger transfer rules. The EU, UK, UAE, KSA and others restrict cross-border transfers of personal data unless safeguards apply. Some Gulf clients, especially in regulated sectors, may require data to stay in-region.

Sensitive data. Health, religion, ethnicity, biometric data (including voiceprints and face geometry used for identification), children's data and financial data attract stricter rules. Avoid putting them into AI tools unless you have a clear legal basis, strong safeguards and, often, explicit consent.

Profiling and targeting. Using AI to infer sensitive characteristics, such as guessing religion from names or health from purchases, for targeting is high risk legally and ethically, and often banned by ad platforms.

Transparency. Update your privacy notice to mention AI tools where they process personal data: what, why and which kinds of providers.

A data protection impact assessment (DPIA), simplified

For higher-risk AI uses (large-scale profiling, sensitive data, new technologies affecting people significantly), many laws expect a DPIA. A simple version asks:

  1. What personal data, whose, and for what purpose?
  2. What is the legal basis?
  3. Which tools and providers process it, and where?
  4. What are the risks to individuals?
  5. What safeguards reduce those risks (minimization, human review, security, retention)?
  6. Is the remaining risk acceptable?

Worked example

A Riyadh e-commerce brand wants to use AI to analyze customer service chats to improve products. Their approach: they use a business-tier AI tool with a data processing agreement and no training on inputs; check data residency preferences with legal advisers given the PDPL's transfer rules; strip names, phone numbers and order IDs before analysis; update their privacy notice; keep outputs aggregated as themes, not individual profiles; and delete raw exports after 30 days.

Hands-on: redact before you prompt

The cheapest data protection control is not sending personal data at all. This script masks common identifiers in exported chats or reviews before you paste them into an AI tool. It is a first pass, not a guarantee: names in free text, addresses and unusual formats still need a human skim.

import re
import sys

PATTERNS = [
    ("EMAIL", re.compile(r"[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}")),
    # International numbers incl. Pakistan (+92), UAE (+971), KSA (+966), UK (+44), US (+1)
    ("PHONE", re.compile(r"(?:\+|00)\d{1,3}[\s-]?\(?\d{1,4}\)?(?:[\s-]?\d{2,4}){2,4}")),
    # Local mobile formats such as 0300-1234567, 050 123 4567, 07700 900123
    ("PHONE", re.compile(r"\b0\d{2,4}[\s-]?\d{3,4}[\s-]?\d{3,4}\b")),
    ("CARD", re.compile(r"\b(?:\d[ -]?){13,19}\b")),
    ("ORDER_ID", re.compile(r"\b(?:ORD|INV|#)[-\s]?\d{4,}\b", re.IGNORECASE)),
]


def redact(text: str) -> tuple[str, dict]:
    counts = {}
    for label, pattern in PATTERNS:
        text, n = pattern.subn(f"[{label}]", text)
        counts[label] = counts.get(label, 0) + n
    return text, counts


if __name__ == "__main__":
    if len(sys.argv) != 3:
        sys.exit("usage: python redact.py input.txt output.txt")
    try:
        with open(sys.argv[1], encoding="utf-8") as f:
            original = f.read()
    except OSError as err:
        sys.exit(f"could not read input: {err}")
    cleaned, counts = redact(original)
    with open(sys.argv[2], "w", encoding="utf-8") as f:
        f.write(cleaned)
    print("masked:", counts)

Hands-on: a one-page DPIA record for an AI use

AI USE:        Summarize customer service chats into monthly product themes
DATA:          Chat text (may include names, phones, order IDs); ~4,000 chats/month; customers in UAE and KSA
PURPOSE:       Product improvement (not marketing profiles)
LEGAL BASIS:   [legitimate interests / other basis recognized in each law] - confirm with adviser
TOOL:          [vendor, business plan], DPA signed [date], no training on inputs, data region [x]
TRANSFERS:     KSA data leaving the Kingdom? -> assess under PDPL transfer rules
RISKS:         Exposure of identifiers; inference about individuals; retention in vendor logs
SAFEGUARDS:    Redaction script before upload; outputs aggregated only; access limited to 3 staff;
               raw exports deleted after 30 days; vendor retention setting checked
RESIDUAL RISK: Low - approved by [name], review date [date]
PRIVACY NOTICE UPDATED: yes, section "How we use AI tools" [date]

Region notes (check current status)

  • UAE: Federal Decree-Law No. 45 of 2021 (PDPL); DIFC and ADGM have their own data protection laws, and the DIFC regime has specific rules for autonomous and semi-autonomous systems.
  • Saudi Arabia: PDPL enforced since September 2024, with regulations on transfers outside the Kingdom.
  • UK: UK GDPR as amended by the Data (Use and Access) Act 2025, which relaxes some automated decision-making rules while keeping safeguards; changes are commencing in stages.
  • EU: GDPR applies alongside the AI Act; they do not replace each other.
  • Pakistan: no comprehensive data protection law had been enacted as of 2026; apply the same principles anyway, because clients and partner laws will expect it.

Pitfalls

  • "It's public on Instagram, so we can do anything with it." Public availability does not remove data protection.
  • Using consumer AI accounts for customer data.
  • Forgetting employees' and creators' data is personal data too.

Key takeaways

  • Data protection applies whenever AI processes information about identifiable people.
  • Core principles: lawfulness, purpose limitation, minimization, accuracy, storage limitation, security and accountability.
  • Use business tools with processing agreements, manage cross-border transfers and avoid sensitive data.
  • Update privacy notices, assess higher-risk uses with a DPIA and keep humans in significant decisions.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which principle is applied by stripping names and phone numbers before AI analysis of customer chats?
  2. Why might a business plan from an AI provider be required for client data?
  3. A follower's details are publicly visible on Instagram. Does data protection still apply if you process them with AI?

Put it into practice

Map one AI use in your business against the seven principles and complete the six-question DPIA for it.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.