Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APIProduction readiness and capstone · Lesson 17 of 19

Security, data privacy and the production checklist

Article · 16 min · 9 min lecture

Video lecture

Security, data privacy and the production checklist

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Security, privacy, production

  • What happens to your data
  • Legal context
  • Technical controls
  • The production checklist

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The risk surface of an AI integration

Integrating an AI API adds new data flows (your users' content leaves your systems), new failure modes (hallucinations, refusals, prompt injection) and new costs. A production-ready integration treats these like any other third-party processor plus some AI-specific controls.

Data handling: know what the provider does with your data

Before sending customer or confidential data, confirm for each provider, product and plan:

  • Training use: first-party business APIs from Anthropic, OpenAI and Google state that API/business data is not used to train models by default; consumer apps and free tiers may differ (Google's Gemini API terms distinguish unpaid and paid services). Read the current terms for the exact product you use.
  • Retention: providers typically retain API inputs and outputs for a limited period for abuse monitoring, and offer zero data retention (ZDR) arrangements for eligible customers and endpoints. Some features (stored responses, files, batches, caches) store data longer by design; OpenAI's Responses API stores responses by default unless you set store=False.
  • Location: where data is processed; options such as Anthropic's inference_geo parameter on supported models, regional deployments on cloud platforms, and data-residency programs.
  • Subprocessors and agreements: a data processing agreement (DPA), subprocessor list, security certifications (for example SOC 2 reports, ISO 27001).
  • GDPR / UK GDPR: lawful basis, data minimization, transparency, processor contracts, international transfer mechanisms, data subject rights.
  • Gulf and South Asia: Saudi Arabia's Personal Data Protection Law (PDPL) and its regulations; the UAE's federal PDPL plus free-zone regimes (DIFC, ADGM); sector regulators in banking and health. Pakistan has been developing personal data protection legislation; check the current status.
  • AI-specific rules: the EU AI Act's phased obligations (for example transparency for certain AI interactions and generated content, and stricter duties for high-risk uses); advertising and consumer-protection rules for AI-generated marketing (disclose where required, avoid unsubstantiated claims).

Involve your privacy/legal team early, document decisions, and update privacy notices to describe AI processing.

Technical controls

  1. Minimize: send only what the task needs; strip IDs, emails and phone numbers unless required.
  2. Redact PII before sending where possible (pattern-based plus a named-entity detector for names and addresses), and re-insert after if needed.
  3. Secrets hygiene (lesson 2): backend-only keys, secret managers, rotation, scanning.
  4. Prompt-injection defenses: treat user content, documents and web pages as untrusted; never let model output directly trigger sensitive actions without validation and approval; restrict tools and outbound channels.
  5. Output handling: never render model output as raw HTML or execute it; sanitize Markdown; validate structured outputs.
  6. Abuse prevention: authentication, per-user rate limits and quotas, content moderation for user-facing generation, and pass a stable pseudonymous user identifier where the provider supports it (for example OpenAI's safety_identifier) so abuse can be traced without sending personal data.
  7. Logging with care: log metadata broadly; store full prompts/outputs only where necessary, redacted, access-controlled and with short retention.
  8. Human oversight for consequential decisions (credit, hiring, medical, legal).

Hands-on: a simple redaction layer

import re

PATTERNS = {
    "EMAIL": re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"),
    "PHONE": re.compile(r"(?:\+?\d[\d\s-]{8,}\d)"),                  # PK +92, UAE +971, KSA +966, UK +44, US +1
    "CARD": re.compile(r"\b(?:\d[ -]?){13,19}\b"),
    "PK_CNIC": re.compile(r"\b\d{5}-\d{7}-\d\b"),
    "UAE_EID": re.compile(r"\b784-?\d{4}-?\d{7}-?\d\b"),
}

def redact(text: str) -> tuple[str, dict]:
    mapping, counter = {}, 0
    for label, pat in PATTERNS.items():
        for m in set(pat.findall(text)):
            counter += 1
            token = f"[{label}_{counter}]"
            mapping[token] = m
            text = text.replace(m, token)
    return text, mapping

def restore(text: str, mapping: dict) -> str:
    for token, original in mapping.items():
        text = text.replace(token, original)
    return text

clean, m = redact("Call Aisha on +971 50 123 4567 or aisha@example.ae about Emirates ID 784-1990-1234567-1")
print(clean)   # names still present: add an NER step (e.g., a local model) for person names and addresses

Regexes miss plenty (names, addresses, free-form IDs) and can over-match; treat this as one layer and test it on real samples.

The production checklist

AreaReady when
QualityEval set per task with agreed thresholds; per-language results
ReliabilityTyped error handling, backoff, circuit breaker, fallback tested
CostPer-call ledger, budgets and alerts, cost per outcome known
SecurityBackend-only keys, secret manager, scanning, injection tests, output sanitization
PrivacyDPA and terms reviewed, retention settings configured, redaction in place, privacy notice updated
ObservabilityTraces/logs with request IDs, dashboards, alerts on errors, latency and spend
OperationsModel IDs and prompts in versioned config, rollback tested, owner and runbook
ComplianceDisclosure where required, human oversight for consequential uses, records of decisions

Worked example: a Lahore fintech's support copilot

The copilot drafts replies for agents. Controls: CNIC numbers, card numbers and phone numbers redacted before calls; provider chosen with a signed DPA and ZDR on the relevant endpoint; drafts never sent automatically; agents see a "draft generated with AI" label; logs keep metadata for 90 days and redacted text for 14; quarterly injection tests using real attack patterns from their inbox. The security review passed in one cycle because every checklist item had evidence.

Pitfalls

  • Sending full records when a few fields would do.
  • Assuming "not used for training" means "not stored".
  • Logging raw prompts with personal data indefinitely.
  • Rendering model output as HTML.

Measuring success

Checklist completion with evidence, PII detection rate on test samples, security test pass rate, time to answer a customer data request, and incidents.

Key takeaways

  • Confirm training use, retention, location and agreements for each provider product and plan.
  • Apply GDPR, UK GDPR, KSA PDPL, UAE PDPL and sector rules with your legal team; watch AI-specific duties.
  • Minimize and redact data, keep keys backend-only, defend against injection and sanitize outputs.
  • Log with care: metadata broadly, sensitive content redacted, restricted and short-lived.
  • Use a production checklist with evidence across quality, reliability, cost, security, privacy and operations.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A provider states API data is not used for training. What else must you check?
  2. Which is the safest way to handle model output in a web page?
  3. Your redaction regex misses customer names. What should you do?

Put it into practice

Complete the production checklist for one AI feature with links to evidence, add a redaction layer, and review your provider's current data terms and retention settings.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.