---
title: "Training data safety, licensing and privacy"
description: "Training data is a liability as well as an asset Once data is baked into weights it is hard to remove. A fine-tuned model can memorize and later…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/data-safety-licensing-and-privacy
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Training data: design, synthesis and safety · lesson 5 of 16 · 15 min

# Training data safety, licensing and privacy

## Training data is a liability as well as an asset

Once data is baked into weights it is hard to remove. A fine-tuned model can **memorize** and later reproduce training examples, including personal data or confidential text, especially rare or repeated strings. Deleting a record from your database does not delete it from a model. So the safest training data is data you are entitled to use, minimized to what the task needs, and scrubbed of what it does not. This lesson is practical guidance, not legal advice.

## Four questions for every dataset

1. **Rights:** Are we allowed to use this data for training? Consider privacy law (lawful basis and purpose), customer contracts, platform terms, copyright and database rights, and dataset licenses.
2. **Minimization:** Does the task need personal data at all? An intent classifier does not need names, phone numbers or account numbers.
3. **Security:** Where is training data stored, who can access it, and where does training happen (your GPUs, a cloud region, a hosted API)?
4. **Retention and deletion:** How long do we keep raw data, cleaned data and models, and what happens when someone exercises a deletion right?

## Privacy frameworks you will meet

- **GDPR / UK GDPR:** lawful basis, purpose limitation (training may be a new purpose relative to why data was collected), data minimization, transparency, and data subject rights. Regulators such as the UK ICO and EU data protection authorities have published guidance on AI and personal data.
- **Saudi PDPL and UAE PDPL** (plus DIFC and ADGM regimes): rules on processing, consent or other legal bases, and cross-border transfer that affect where you can train.
- **Sector rules** (health, finance, telecom) often go further.

Where training involves significant personal data or high-impact decisions, conduct a **data protection impact assessment** before you start.

## PII handling in practice

- **Detect and redact** with a tool such as Microsoft Presidio plus custom recognizers for local identifiers (CNIC, Emirates ID, Saudi national ID formats, IBANs).
- **Replace, don't just delete:** swap real values for consistent fake ones ("Ayesha" → "Customer_A", account numbers → a fake format) so the model still learns structure.
- **Check redaction quality** on a sample; recognizers miss things, especially in Arabic, Urdu and mixed scripts.
- **Keep a mapping only if necessary**, stored separately with strict access, or not at all.

```python
# redact.py (pip install presidio-analyzer presidio-anonymizer; also install a spaCy English model per Presidio docs)
from presidio_analyzer import AnalyzerEngine, Pattern, PatternRecognizer
from presidio_anonymizer import AnonymizerEngine
from presidio_anonymizer.entities import OperatorConfig

cnic = PatternRecognizer(supported_entity="PK_CNIC", patterns=[Pattern("cnic", r"\b\d{5}-\d{7}-\d\b", 0.9)])
emirates_id = PatternRecognizer(supported_entity="AE_EID", patterns=[Pattern("eid", r"\b784-\d{4}-\d{7}-\d\b", 0.9)])

analyzer = AnalyzerEngine()
analyzer.registry.add_recognizer(cnic)
analyzer.registry.add_recognizer(emirates_id)
anonymizer = AnonymizerEngine()

def redact(text: str) -> str:
    results = analyzer.analyze(text=text, language="en")
    return anonymizer.anonymize(text=text, analyzer_results=results, operators={
        "DEFAULT": OperatorConfig("replace", {"new_value": "<REDACTED>"}),
        "PERSON": OperatorConfig("replace", {"new_value": "<NAME>"}),
        "PK_CNIC": OperatorConfig("replace", {"new_value": "00000-0000000-0"}),
        "AE_EID": OperatorConfig("replace", {"new_value": "784-0000-0000000-0"}),
    }).text

print(redact("I'm Ayesha Khan, CNIC 42101-1234567-1, email ayesha@example.com"))
```

Test on your own languages; add recognizers and review samples until the miss rate is acceptable.

## Copyright and content licenses

- Customer-generated content, scraped web text, purchased datasets and open datasets all come with terms. **CC BY** requires attribution; **CC BY-SA** adds share-alike; **NC** licenses forbid commercial use.
- Copyright law on AI training differs across jurisdictions and is still evolving through courts and legislation. Prefer data you created, licensed explicitly for training, or that your contracts permit.
- **The base model's license** also applies to your fine-tuned derivative (Module 1 of the open-weight course): attribution, naming and use policies carry through.

## Regulation touchpoints for fine-tuners

Under the EU AI Act, providers of general-purpose AI models have obligations (documentation, copyright policy, training-content summary). Guidance from the European Commission indicates that only substantial modifications (assessed with a compute-based indicator) make a downstream modifier a GPAI model provider; typical small LoRA fine-tunes are far below that level, but check the current guidance for your case. Separately, **your use case** may be high-risk or trigger transparency obligations regardless of how you trained.

## Safety of the tuned model

Fine-tuning can erode a base model's safety behavior, even with benign data. Include **safety evaluations** (refusal of harmful requests, no leakage of training PII, jailbreak resistance) in your before/after tests, and mix in safety examples if needed.

## Worked example: a clinic network in Pakistan

A clinic network wants a model that drafts follow-up reminders from appointment notes. They minimize first: the model needs appointment type and timing, not diagnosis details. They redact names and CNICs with custom recognizers, keep training on their own servers, run an impact assessment, and add a memorization test (prompting with the start of real notes and checking that the model does not complete sensitive details). Patient notices are updated to reflect the use of data for service improvement where the law requires it.

## Pitfalls

- "We anonymized it" when only names were removed (quasi-identifiers remain).
- Training on customer data without checking contracts and notices.
- Forgetting that deletion requests may require retraining or removing a model.
- Skipping safety evaluation after tuning.

## How to measure success

A completed rights review, a data card with PII handling and measured redaction miss rate, an impact assessment where required, and safety plus memorization tests in your evaluation suite.

## Video lecture: Training data safety, licensing and privacy

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Data safety, licensing, privacy
2. Analogy: baking a cake
3. Four questions
4. Privacy frameworks
5. Redaction
6. Copyright and licenses
7. Regulation touchpoints
8. Safety after tuning
9. Worked example: clinic network
10. Simple example: estate agency
11. Document for trust
12. FAQ: train on support transcripts?
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### Data safety, licensing, privacy

Here is an uncomfortable fact. If a training example contains a customer's phone number, the fine-tuned model might one day repeat it to someone else. And deleting that customer from your database does not delete them from the model. In this lesson you will learn to treat training data as both an asset and a liability: rights, minimization, PII redaction, copyright and licenses, regulation touchpoints, and safety testing after tuning. This is practical guidance, not legal advice.

### Analogy: baking a cake

Here is an analogy. Training a model on data is like mixing ingredients into a cake. Before baking, you can still pick out a raisin you did not want. After baking, the raisin is inside, and getting it out means baking a new cake. Personal data in training sets is that raisin. So you check the ingredients before you bake: do we have the right to use them, do we need them at all, and have we removed what should not be there?

### Four questions

Ask four questions of every dataset. Rights: are we allowed to use this for training, under privacy law, contracts, platform terms, copyright and dataset licenses? Minimization: does the task need personal data at all? An intent classifier does not need names or account numbers. Security: where is data stored, who can access it, and where does training run? And retention and deletion: how long do we keep raw data, cleaned data and models, and what happens when someone asks to be deleted?

### Privacy frameworks

Privacy frameworks shape the answers. Under GDPR and UK GDPR, training can be a new purpose compared with why you collected the data, so you need a lawful basis, transparency and respect for data subject rights. Regulators, like the UK ICO, have published guidance on AI and personal data. Saudi Arabia's and the UAE's data protection laws, plus DIFC and ADGM regimes, affect processing and where training can happen. Sector rules often go further. For significant personal data or high-impact uses, do an impact assessment first.

### Redaction

Now redaction. Detect personal data with a tool like Microsoft Presidio, and add custom recognizers for local identifiers, like Pakistani CNIC numbers or Emirates ID numbers. Replace rather than delete: swap real values for consistent fake ones so the model still learns the structure of a message. Then check quality on a sample, because recognizers miss things, especially in Arabic, Urdu and mixed scripts. Keep a mapping back to real values only if you truly need it, locked away separately.

### Copyright and licenses

Copyright and licenses next. Customer content, scraped text, bought datasets and open datasets all have terms. CC BY needs attribution, share-alike adds conditions, and non-commercial licenses forbid commercial use. Copyright law on AI training differs by country and is still evolving, so prefer data you created, licensed for training, or that your contracts permit. And remember the base model's license carries into your fine-tuned version.

### Regulation touchpoints

Regulation touchpoints. Under the EU AI Act, providers of general-purpose AI models have documentation and copyright duties. European Commission guidance says only substantial modifications, judged with a compute-based indicator, turn someone who modifies a model into a provider of a general-purpose model. Typical small LoRA fine-tunes are far below that, but check the current guidance. Separately, your use case may be high-risk or carry transparency duties regardless of how you trained.

### Safety after tuning

Fine-tuning can also weaken a model's safety behavior, even with harmless data. So add safety evaluations to your before-and-after tests: refusals of harmful requests, resistance to jailbreak attempts, and no leakage of training data. A simple memorization test prompts the model with the start of real training records and checks that it does not complete sensitive details. If safety drops, mix safety examples into training.

### Worked example: clinic network

A worked example. A clinic network in Pakistan wants a model that drafts follow-up reminders from appointment notes. They minimize first: the model needs the appointment type and timing, not diagnoses. They redact names and CNICs with custom recognizers, train on their own servers, run an impact assessment, add a memorization test, and update patient notices where required. The model is useful, and the risk is proportionate.

### Simple example: estate agency

A simple example. An estate agency wants a model that classifies inquiry emails. The emails contain names, phone numbers and addresses. For classification, none of that is needed. So they replace names with Customer, phone numbers with a fixed fake number, and addresses with City, before training. They check fifty redacted emails by hand, find three missed phone numbers in an unusual format, add a pattern for it, and re-run. The model learns the intent, not the people.

### Document for trust

Document it so others can trust it. Your data card should record the lawful basis or contractual permission you rely on, what personal data was removed and how, the measured redaction miss rate, the licenses of any external datasets, the base model license, and who approved the use. When a regulator, customer or auditor asks how the model was trained, you want to hand over a document, not start an investigation.

### FAQ: train on support transcripts?

A practical question: can we use customer support transcripts to train a model? Possibly, but check three things first. Does your privacy notice, contract or other legal basis cover this use? Can you remove or replace personal data without losing what the model needs to learn? And do any customer contracts restrict how their data may be used? If all three check out, document the decision in the data card, and consider giving enterprise customers a way to opt out.

### Try this now

Try this now. Take ten real training examples and highlight every piece of personal data by hand. Then run an automatic detector on the same ten and compare. Count what the detector missed. That miss count, even on ten examples, tells you whether your redaction is ready or needs custom patterns first.

### Watch me do it

Watch me do it. I take one hundred training examples and run the redaction script with the default recognizers plus our CNIC and Emirates ID patterns. Then I review every redacted example by hand, with the original beside it. I find six misses: four phone numbers written with spaces, one name in Urdu script, one account number inside a sentence. I add a phone pattern that accepts spaces and an account-number pattern, rerun, and review again: one miss left, the Urdu name. I note it as a known limitation and add a human review step for Urdu-script messages. Then the rights check: I open our privacy notice and the relevant customer contract clause, and write in the data card the basis we rely on and who approved it. Finally I add a memorization test to the evaluation plan: ten prompts starting with the beginning of real records, checking the model does not complete sensitive details.

### Recap

Recap. Models can memorize, so treat training data as a liability. Ask about rights, minimization, security and deletion. Redact and replace personal data, and measure what you miss. Respect copyright and licenses. Know your regulation touchpoints. And re-test safety after tuning. Your next step: run the redaction script on a hundred of your examples, review misses by hand, and record the miss rate in your data card.

## Key takeaways

- Fine-tuned models can memorize training data; removal after training is hard
- Ask four questions: rights, minimization, security, retention/deletion
- Redact and replace PII with tested recognizers, including local ID formats
- Check copyright, dataset licenses and the base model license
- Re-run safety and memorization tests after fine-tuning

## Try it

Run the redaction script (with local recognizers) on 100 of your training examples, review misses manually, and record the miss rate in your data card.

- [Previous: Synthetic data generation and distillation](https://optimizeall.com/learn/fine-tuning-and-custom-models/synthetic-data-and-distillation)
- [Next: Supervised fine-tuning fundamentals](https://optimizeall.com/learn/fine-tuning-and-custom-models/sft-fundamentals)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
