---
title: "Governing AI: human review, audit trails and data quality"
description: "Why governance is non-negotiable Project controls outputs influence funding decisions, contractual positions, claims and public accountability. If an…"
url: https://optimizeall.com/learn/project-controls-with-ai/ai-governance
updated: 2026-10-05
---

Project Controls in the AI Era · AI in project controls: use cases and governance · lesson 21 of 22 · 15 min

# Governing AI: human review, audit trails and data quality

## Why governance is non-negotiable

Project controls outputs influence funding decisions, contractual positions, claims and public accountability. If an AI-assisted forecast is wrong, someone must be able to explain what happened, who reviewed it and on what basis it was accepted. Governance makes AI use **safe, explainable and auditable**.

## Principles

1. **Human accountability.** A named person owns every forecast, report and decision. AI assists; it does not sign.
2. **Human review proportional to risk.** The higher the impact, the more rigorous the review.
3. **Traceability.** Inputs, prompts or model versions, outputs and human edits are recorded.
4. **Data quality and fitness.** Models use governed, validated data.
5. **Confidentiality and privacy.** Only approved tools; sensitive data protected; personal data handled lawfully.
6. **Transparency.** Readers know when content was AI-assisted, per company policy.
7. **Monitoring.** Model performance is tracked over time and models are retired or retrained when they drift.

## A risk-tiered review model

| Tier | Example | Review requirement |
|---|---|---|
| Low | Formatting, summarising meeting notes | Spot check |
| Medium | Draft variance narrative, schedule quality suggestions | Full review by controls lead; facts checked against source |
| High | EAC or forecast date used for funding, contractual notices, claims | Independent review, documented rationale, approval by accountable manager |
| Prohibited | Fully automated contractual or financial decisions without human approval | Not allowed |

## The audit trail

For each AI-assisted output, record:

```
Output ID: RPT-2026-09-P14-narrative
Purpose: Monthly executive narrative
Tool / model and version: [approved enterprise assistant, version]
Input data: EVM table v3 (data date 31-Aug), risk register export 01-Sep, change log 01-Sep
Prompt / template reference: NARR-TPL-02
AI draft stored: yes (link)
Human reviewer: [name]   Date: ______
Changes made: corrected cause for CA 1.3; removed unsupported claim on recovery
Approval: [accountable manager]
```

This record allows anyone to reconstruct how an output was produced and reviewed, which is increasingly expected by auditors and, in some sectors and jurisdictions, by regulation.

## Data quality for AI

AI-specific checks on top of normal controls data quality:

- **Representativeness:** is training history similar to the current project (type, size, region, contract model)?
- **Leakage:** does the model accidentally use information that would not have been known at the time (e.g., final costs) when trained on history?
- **Bias:** are some project types or contractors under-represented, leading to systematically wrong predictions?
- **Freshness:** are prices, productivity and conditions in the data still relevant?

## Validating models

- **Back-testing:** run the model on completed projects as they were at earlier data dates and compare its forecasts with actual outcomes, alongside traditional methods.
- **Benchmark:** a model should beat simple methods (e.g., CPI-based EAC) to justify its complexity.
- **Explainability:** prefer outputs that show drivers ("this forecast is higher because CPI has declined for four periods and similar packages overran").
- **Drift monitoring:** track error over time; investigate if it grows.

## Regulatory and policy landscape (high level)

Rules on AI vary by jurisdiction and continue to evolve. The EU AI Act introduces risk-based obligations; the UK has taken a principles-based, regulator-led approach; the US has a mix of federal guidance and state laws; the UAE and Saudi Arabia have published national AI strategies, principles and guidance, and Pakistan has developed national AI policy work. Management-system standards such as ISO/IEC 42001 provide a framework for AI governance. Always follow your organisation's legal and compliance guidance; this course does not provide legal advice.

## Worked example: a governance failure

*Illustrative.* A fictional developer used an AI model to forecast EAC for a portfolio. The model had been trained on projects that had mostly finished, but the training data inadvertently included final-cost fields, so historical back-tests looked excellent. In live use, forecasts were far too optimistic. Because no audit trail recorded which model version produced which forecast, it took weeks to identify the problem. Fixes: leakage checks, versioned models, back-testing at historical data dates, and a rule that portfolio forecasts are always compared with CPI-based methods.

## Common mistakes

- "The AI said so" as a justification.
- No record of model versions or prompts.
- Using public tools with confidential commercial data.
- Training on data that is not representative of current projects.
- No ongoing monitoring after deployment.

## Hands-on: back-testing a forecast method against the CPI benchmark

```python
import pandas as pd

# one row per completed project per historical snapshot, using ONLY data known at that date
snap = pd.read_csv("historical_snapshots.csv")  # project, pct_complete, BAC, EV, AC, model_eac, final_cost
snap["cpi_eac"] = snap.BAC / (snap.EV / snap.AC)
for m in ["model_eac", "cpi_eac"]:
    snap[f"err_{m}"] = (snap[m] - snap.final_cost).abs() / snap.final_cost

summary = snap.groupby("pct_complete")[["err_model_eac", "err_cpi_eac"]].median().round(3)
print(summary)            # median absolute % error by stage
wins = (snap.err_model_eac < snap.err_cpi_eac).mean()
print(f"Model beats CPI benchmark in {wins:.0%} of snapshots")
```

Leakage check before training: list every feature and ask "would this have been known at the snapshot date?" Remove anything that encodes the outcome (final cost, final duration, close-out flags).

## Template: AI use policy for a controls team (one page)

```text
1 Purpose and scope: which controls outputs may be AI-assisted.
2 Approved tools: names, versions, data classifications each may process.
3 Tiers: Low (spot check) | Medium (full review vs source) | High (independent review + approval) | Prohibited.
4 Required audit fields: output ID, purpose, tool/model version, inputs + data dates, prompt/template ref,
  draft stored, reviewer, changes made, approver.
5 Data rules: no confidential commercial or personal data in unapproved tools; leakage and bias checks for models.
6 Validation and monitoring: back-test vs simple benchmarks; quarterly error review; retire on drift.
7 Transparency: how AI assistance is disclosed in reports (per company policy).
8 Ownership: policy owner, review date.
```

## How to measure success

- 100% of medium and high-tier outputs have a complete audit record.
- Models in use beat simple benchmarks in back-tests and are re-checked at least quarterly.
- No incidents of unapproved tools processing confidential data.

## Video lecture: Governing AI: human review, audit trails and data quality

Lecture coming soon · 9 chapters · about 8 minutes. Read the full transcript below.

1. 'The AI said so' is not an answer
2. Why it matters
3. The concept: seven principles
4. Risk-tiered review
5. Worked example one: an audit trail entry
6. Worked example two: a leakage failure
7. Watch me do it: validating a forecast model
8. Regulation and standards, briefly
9. Common mistakes, recap and try this now

## Lecture transcript

### 'The AI said so' is not an answer

Imagine a steering committee has released funding based on a forecast. Six months later, it's badly wrong. Someone asks: how was this number produced, who checked it and on what basis was it accepted? If the answer is 'the AI said so', you have a governance failure, whatever the model's accuracy. In project controls, outputs influence funding, contracts, claims and sometimes public accountability. So AI use has to be safe, explainable and auditable. In this lecture you'll learn seven principles for governing AI in controls, a risk-tiered review model, what an audit trail should record, AI-specific data quality checks including leakage, and how to validate a model before trusting it. By the end, you'll be able to draft a one-page AI use policy for your team.

### Why it matters

Why does governance matter more in controls than in many other uses of AI? Because the outputs have consequences. A forecast date can trigger a contractual notice. An EAC can release or withhold millions. A risk assessment can shape a claim. When those outputs are wrong, someone must be able to reconstruct what happened. And the expectation of that record is growing. Auditors increasingly ask how AI-assisted outputs were produced and reviewed, and in some sectors and jurisdictions, regulation adds formal obligations. Good governance isn't there to slow you down. It's what lets you use AI in serious work at all, because it gives leaders a reason to trust the result.

### The concept: seven principles

Here are seven principles. One, human accountability: a named person owns every forecast, report and decision. AI assists. It doesn't sign. Two, review proportional to risk: the higher the impact, the more rigorous the review. Three, traceability: inputs, prompts or model versions, outputs and human edits are recorded. Four, data quality and fitness: models use governed, validated data. Five, confidentiality and privacy: only approved tools, sensitive data protected, personal data handled lawfully. Six, transparency: readers know when content was AI-assisted, according to company policy. And seven, monitoring: model performance is tracked over time, and models are retrained or retired when they drift. Think of it like the controls you already apply to a junior analyst's work, made explicit and written down.

### Risk-tiered review

Now the review model, with four tiers. Low: formatting, or summarising meeting notes. A spot check is enough. Medium: a draft variance narrative, or schedule quality suggestions. The controls lead does a full review and checks facts against the source. High: an EAC or a forecast date used for funding decisions, contractual notices or claims. That needs an independent review, a documented rationale and approval by the accountable manager. And prohibited: fully automated contractual or financial decisions without human approval. Here's the key idea. The tier depends on what the output is used for, not on how clever the tool is. The same model can produce a low-tier summary and a high-tier forecast. The review follows the use.

### Worked example one: an audit trail entry

Let's look at a simple audit trail entry. Output ID: the September monthly executive narrative for project fourteen. Purpose: monthly executive narrative. Tool and version: the approved enterprise assistant, with its version. Input data: EVM table version three, data date the thirty-first of August; the risk register export and change log from the first of September. Prompt: the standard narrative template, reference NARR two. AI draft stored: yes, with a link. Human reviewer and date. Changes made: corrected the cause attributed to control account one point three, and removed an unsupported claim about recovery. And approval by the accountable manager. It takes two minutes to complete. And if anyone asks, a year later, how that narrative was produced, the answer is on one page.

### Worked example two: a leakage failure

Now a realistic failure, from the lesson. A fictional developer built a model to forecast EAC across its portfolio. Back-tests looked excellent. Almost too good. In live use, forecasts were far too optimistic. The problem? The training data had inadvertently included final-cost fields, information that wouldn't have been known at the time of each forecast. That's called leakage, and it makes a model look brilliant on history and useless in practice. Worse, because no audit trail recorded which model version produced which forecast, it took weeks to work out what had gone wrong. The fixes: explicit leakage checks, versioned models, back-testing using snapshots as they were at historical data dates, and a rule that portfolio forecasts are always compared with simple CPI-based methods.

### Watch me do it: validating a forecast model

Here's how I'd validate a forecasting model before anyone relies on it. I take completed projects and use the snapshots stored at historical data dates: say, at twenty-five, fifty and seventy-five per cent complete. At each snapshot, I compute two forecasts using only information available at that date. The model's forecast, and a simple benchmark: budget divided by CPI. Then I compare each with the final actual cost and calculate the error. If the model doesn't clearly beat the simple benchmark, it doesn't justify its complexity. If it does, I deploy it as a third opinion in forecast reviews, alongside formulas and the bottom-up estimate. And I keep monitoring its error on new projects, because models drift as prices, productivity and contract types change.

### Regulation and standards, briefly

A brief word on regulation, which varies by jurisdiction and keeps evolving. The EU AI Act takes a risk-based approach, with obligations phasing in over several years, so check the current timelines for your use case. The UK has taken a principles-based approach led by existing regulators. The US has a mix of federal guidance and state laws. The UAE and Saudi Arabia have published national AI strategies, principles and guidance, and Pakistan has developed national AI policy work. On the management side, ISO slash IEC forty-two thousand and one sets out requirements for an AI management system, and frameworks such as the NIST AI Risk Management Framework are widely used. Always follow your organisation's legal and compliance guidance. This course doesn't provide legal advice.

### Common mistakes, recap and try this now

The common mistakes: 'the AI said so' as a justification, no record of model versions or prompts, using public tools with confidential commercial data, training on data that isn't representative of current projects, and no monitoring after deployment. Also check for bias: if some project types or contractors are under-represented in the history, predictions for them can be systematically wrong. So, to recap. Apply seven principles, led by human accountability. Tier your review by how the output will be used. Keep an audit trail for every AI-assisted output. Check data for representativeness, leakage, bias and freshness. And validate models by back-testing against simple benchmarks. Your try-this-now: draft a one-page AI use policy for your controls team, covering tiers, review requirements, approved tools, audit trail fields and prohibited uses.

## Key takeaways

- A named human is accountable for every AI-assisted forecast, report and decision.
- Scale review rigour to impact; prohibit fully automated high-stakes decisions.
- Record inputs, model/version, outputs, human edits and approvals in an audit trail.
- Validate models with back-testing, benchmarks, leakage and bias checks, and monitor drift.

## Try it

Draft a one-page AI use policy for your controls team: tiers, review requirements, approved tools, audit trail fields and prohibited uses.

- [Previous: Where AI adds value in project controls](https://optimizeall.com/learn/project-controls-with-ai/ai-use-cases)
- [Next: An AI adoption roadmap for controls teams](https://optimizeall.com/learn/project-controls-with-ai/ai-adoption-roadmap)
- [All lessons of Project Controls in the AI Era](https://optimizeall.com/learn/project-controls-with-ai)
