---
title: "Lab: an EVM and Monte Carlo notebook in Python"
description: "What you will build In this lab you build a small, reusable Jupyter notebook that does the analytical core of a monthly controls cycle: earned value…"
url: https://optimizeall.com/learn/project-controls-with-ai/evm-and-monte-carlo-python-lab
updated: 2026-10-05
---

Project Controls in the AI Era · Risk management, Monte Carlo and change control · lesson 15 of 22 · 25 min

# Lab: an EVM and Monte Carlo notebook in Python

## What you will build

In this lab you build a small, reusable Jupyter notebook that does the analytical core of a monthly controls cycle: earned value metrics and trends by control account, an EAC range with TCPI, a simple anomaly flag, and a Monte Carlo schedule risk analysis with a discrete risk, correlation, P50/P80, criticality and a tornado ranking. It ends by writing a structured status table that a controls lead reviews and that an approved AI assistant can use to draft a narrative. Every number is illustrative; the method is what you reuse.

The notebook deliberately stays small so you can read every line. Enterprise tools (Primavera P6, Microsoft Project, Power BI, Safran Risk, Deltek Acumen Risk, @RISK) do all of this at scale; building it once by hand is how you learn to check them.

## Setup

```bash
python -m venv .venv && source .venv/bin/activate      # Windows: .venv\Scripts\activate
pip install pandas numpy jupyter matplotlib
jupyter notebook
```

Create a notebook called `controls_lab.ipynb` and paste each cell below into its own cell.

## Cell 0 and 1: imports and data

In real use, cell 1 is a `pd.read_csv()` of your monthly export by control account and period (cumulative PV, EV and AC, plus BAC). Keep the shape: one row per account per period.

```python
import numpy as np
import pandas as pd

# Cell 1: cumulative EVM data by control account and period (USD 000, illustrative)
data = pd.DataFrame({
    "account": ["Civil"] * 4 + ["Pumps"] * 4,
    "period":  [1, 2, 3, 4] * 2,
    "BAC": [2400] * 4 + [5000] * 4,
    "PV":  [300, 800, 1300, 1900, 400, 1100, 1900, 2800],
    "EV":  [290, 760, 1250, 1800, 380, 950, 1550, 2200],
    "AC":  [295, 780, 1280, 1850, 400, 1080, 1850, 2700],
})
```

## Cell 2: indices, variances and the EAC range

Notice that every metric is computed from cumulative values, and that indices are never averaged across accounts.

```python
# Cell 2: indices, variances and EAC range per account and period
d = data.copy()
d["CPI"], d["SPI"] = d.EV / d.AC, d.EV / d.PV
d["CV"], d["SV"] = d.EV - d.AC, d.EV - d.PV
d["EAC_cpi"] = d.BAC / d.CPI
d["EAC_comp"] = d.AC + (d.BAC - d.EV) / (d.CPI * d.SPI)
d["TCPI_bac"] = (d.BAC - d.EV) / (d.BAC - d.AC)
latest = d[d.period == d.period.max()]
print(latest[["account", "CPI", "SPI", "EAC_cpi", "EAC_comp", "TCPI_bac"]].round(2))
```

Reading the output: Pumps has CPI 0.81 and SPI 0.79. Its CPI-method EAC is about 6.14M against a 5.0M budget, the composite method gives about 7.07M, and the TCPI to BAC is 1.22. A budget-level forecast for Pumps is not credible without a specific change. Civil is close to plan.

## Cell 3: a simple anomaly flag

Anomaly detection does not have to start with machine learning. A rule such as "CPI has fallen for three consecutive periods" catches a large share of what matters and is easy to explain.

```python
# Cell 3: simple anomaly flag - CPI falling for 3 consecutive periods
d["dCPI"] = d.groupby("account")["CPI"].diff()
falling = d.groupby("account")["dCPI"].apply(lambda s: (s.tail(3) < 0).all())
print("CPI falling 3 periods:", list(falling[falling].index))
```

Later you can add statistical flags (for example a z-score on period CPI or on cost postings) or a model; keep the simple rules as a baseline the model must beat.

## Cell 4: Monte Carlo schedule risk

Five activities, PERT distributions from three-point estimates, a shared productivity factor that correlates the two site activities, and one discrete risk from the register.

```python
# Cell 4: schedule Monte Carlo on a small network (days)
rng = np.random.default_rng(7)
N = 20_000

def pert(a, m, b, size, lam=4):
    """PERT (beta) sample from min a, most likely m, max b."""
    alpha = 1 + lam * (m - a) / (b - a)
    beta_ = 1 + lam * (b - m) / (b - a)
    return a + rng.beta(alpha, beta_, size) * (b - a)

acts = {  # id: (min, most likely, max, predecessors)
    "Design":      (18, 20, 30, []),
    "Procure":     (40, 45, 70, ["Design"]),
    "Civil":       (30, 35, 50, ["Design"]),
    "Install":     (25, 28, 40, ["Procure", "Civil"]),
    "Commission":  (8, 10, 16, ["Install"]),
}
# correlation via a shared productivity factor on site work (Civil, Install)
site_factor = rng.normal(1.0, 0.08, N).clip(0.8, 1.3)
dur = {a: pert(*acts[a][:3], N) for a in acts}
dur["Civil"] *= site_factor
dur["Install"] *= site_factor
# discrete risk from the register: R-07 grid approval, 30% chance, +10 to +25 days on Commission
r07 = rng.random(N) < 0.30
dur["Commission"] += np.where(r07, rng.uniform(10, 25, N), 0)

ef = {}
for a, (_, _, _, preds) in acts.items():  # dict order is topological here
    es = np.maximum.reduce([ef[p] for p in preds]) if preds else np.zeros(N)
    ef[a] = es + dur[a]
finish = ef["Commission"]
plan = 20 + 45 + 28 + 10  # deterministic most-likely critical path
print(f"Plan {plan} d | confidence {np.mean(finish <= plan):.0%} | "
      f"P50 {np.percentile(finish, 50):.0f} | P80 {np.percentile(finish, 80):.0f}")
```

Typical output: a deterministic plan of 103 days with only about 11% confidence, P50 around 113 days and P80 around 125 days. Two effects drive the gap: right-skewed ranges (maximums further from the most likely than minimums) and the 30% risk event on commissioning.

## Cell 5: criticality and a tornado ranking

```python
# Cell 5: criticality of the Procure vs Civil branch, and a simple tornado (rank correlation)
crit_procure = np.mean(ef["Procure"] >= ef["Civil"])
print(f"Procure drives Install in {crit_procure:.0%} of runs; Civil in {1 - crit_procure:.0%}")
rank = lambda x: pd.Series(x).rank().to_numpy()
tornado = {a: np.corrcoef(rank(dur[a]), rank(finish))[0, 1] for a in acts}
for a, r in sorted(tornado.items(), key=lambda kv: -abs(kv[1])):
    print(f"{a:11s} {r:+.2f}")
```

Procurement drives installation in almost every iteration, so it is effectively critical; commissioning ranks highest in the tornado because it carries the discrete risk. That tells you where mitigation buys the most confidence: the grid-approval risk and procurement lead time.

## Cell 6: a structured status table

```python
# Cell 6: a structured status table for review (and for an approved AI tool to draft from)
status = latest[["account", "CPI", "SPI", "CV", "EAC_cpi", "TCPI_bac"]].round(2)
status["schedule_P50_days"] = round(float(np.percentile(finish, 50)))
status["schedule_P80_days"] = round(float(np.percentile(finish, 80)))
status.to_csv("status_for_review.csv", index=False)
print(status.to_string(index=False))
```

This CSV is the single, reviewed source for the monthly narrative. If you use an approved enterprise AI assistant to draft the narrative, give it this table (not raw emails or commercial documents) with a template prompt such as:

```text
Draft a 150-word project status narrative from the table below.
Structure: headline; cost (CPI, EAC range, TCPI); schedule (P50/P80 vs plan); top driver; decisions needed.
Cite the column for every number. Do not state causes that are not in the table; write "cause to be confirmed".
```

The controls lead checks every figure against the notebook output, confirms causes with control account managers and records the edits in the audit trail.

## From lab to real projects

| Lab step | Real-world source | Watch out for |
|---|---|---|
| EVM data | ERP and schedule exports by control account and period | Accruals, consistent data dates, codes that match the WBS |
| Network | Primavera P6 or Microsoft Project export (activities, links, durations) | Only finish-to-start handled here; lags and SS/FF links need care |
| Ranges | Structured interviews plus historical actual-vs-plan | Blanket ±10% ranges; anchoring on the plan |
| Risks | Register entries mapped to activities | Double counting with activity ranges |
| Correlation | Shared drivers (productivity, weather, a common supplier) | Ignoring it narrows the range unrealistically |

## Common mistakes

- Averaging CPIs across accounts instead of computing from summed EV and AC.
- Running Monte Carlo on a network with missing logic.
- Reporting P80 to one decimal place of a day, implying false precision.
- Letting an AI-drafted narrative go out without checking every number against this output.

## How to measure success

- The notebook reproduces your scheduling tool's deterministic finish for the same network (a basic validation).
- Results are stable when you change the random seed (use at least 10,000 iterations).
- Each month's status narrative can be traced, number by number, to the status table.

## Video lecture: Lab: an EVM and Monte Carlo notebook you can reuse

Lecture coming soon · 10 chapters · about 9 minutes. Read the full transcript below.

1. Build it once, check every tool forever
2. Why a notebook?
3. The setup
4. Cell two: earned value, done correctly
5. Cell three: a simple anomaly flag
6. Cell four: Monte Carlo schedule risk
7. Watch me do it: reading the results
8. Cell six: a structured status table, and AI
9. Common mistakes and extending the lab
10. Recap and try this now

## Lecture transcript

### Build it once, check every tool forever

Here's something that changed how I work. The first time I built an earned value calculation and a Monte Carlo simulation myself, line by line, I stopped treating my tools as black boxes. I could check them. And when a number in a report looked odd, I could reproduce it in minutes. In this lab lecture I'm going to walk you through a small Jupyter notebook that does the analytical core of a monthly controls cycle. Earned value metrics and trends by control account. An EAC range with the to-complete index. A simple anomaly flag. A Monte Carlo schedule risk analysis with a discrete risk and correlation. And finally, a structured status table for review. By the end you'll have a notebook you can adapt to your own project data.

### Why a notebook?

Why a notebook rather than just using Primavera P6, Power BI or a risk tool? You should use those. They do this at scale, with proper data connections, security and audit. But a notebook gives you three things. Transparency: every calculation is visible and readable, which is exactly what you need when someone challenges a number. Repeatability: the same steps run the same way every month, with no copy-paste errors. And validation: you can check that your enterprise tool is doing what you think it is. If your notebook and your scheduling tool disagree about a deterministic finish date on the same network, one of them has a hidden constraint, a calendar issue or a lag that somebody needs to explain.

### The setup

The setup takes five minutes. Create a Python virtual environment, install pandas, NumPy, Jupyter and matplotlib, and open a new notebook. The exact commands are in the lesson text. Now, the most important design decision is the shape of your data. One row per control account per period, with cumulative planned value, earned value and actual cost, plus the budget at completion. In the lab, I type a small illustrative table straight into the first cell: two accounts, civil and pumps, over four months. In real use, that cell becomes a single line that reads a CSV exported from your ERP and scheduling system. Keep the shape identical and everything downstream just works.

### Cell two: earned value, done correctly

Cell two calculates the metrics. For every account and period: CPI is earned value over actual cost, SPI is earned value over planned value, cost and schedule variances are earned value minus actual cost and earned value minus planned value. Then the CPI-method EAC, the composite EAC and the to-complete index to budget. Look at the latest period. Civil has a CPI of about nought point nine seven. Healthy enough. Pumps has a CPI of nought point eight one and an SPI of nought point seven nine. Its CPI-method EAC is about six point one four million against a five million budget, and the TCPI to budget is one point two two. Here's the key idea: a budget-level forecast for pumps is not credible without a specific, evidenced change in how the work is being done.

### Cell three: a simple anomaly flag

Cell three is an anomaly flag, and I want to make a point about AI here. You don't have to start with machine learning. A simple rule, such as 'CPI has fallen for three consecutive periods', catches a surprising share of what matters, and anyone can understand why it fired. In our data, it flags pumps. Later, you might add a statistical test on cost postings, or a trained model that looks for miscoded costs or duplicate invoices. That's worthwhile. But keep the simple rules running alongside, as a baseline. If a sophisticated model can't beat a transparent rule on your historical data, you don't need the model yet.

### Cell four: Monte Carlo schedule risk

Cell four is the Monte Carlo. There are five activities: design, procure, civil, install and commission. Procurement and civil works run in parallel after design, and installation needs both. Each activity has a minimum, most likely and maximum duration, sampled from a PERT distribution, which gives more weight to the most likely value than a triangular does. Then two refinements that make the model realistic. First, correlation: civil and installation share a site productivity factor, so if productivity is poor, both suffer together. Ignore that and your range comes out unrealistically narrow. Second, a discrete risk from the register, R-07, grid approval: a thirty per cent chance of adding ten to twenty-five days to commissioning. We run twenty thousand iterations, and the finish is simply the early finish of commissioning in each run.

### Watch me do it: reading the results

Let me run it and read the output with you. The deterministic plan, adding the most likely durations along the longest path, is one hundred and three days. The simulation says there's only about an eleven per cent chance of finishing by then. P50 is around one hundred and thirteen days, and P80 around one hundred and twenty-five. Why such a gap? Two reasons. The ranges are skewed: things can go much more wrong than right. And the thirty per cent risk on commissioning adds a long tail. Next, criticality. Procurement drives installation in about ninety-six per cent of runs, so it's effectively critical. And the tornado, which ranks each activity's correlation with the finish, puts commissioning first, because it carries the risk, and procurement second. That's exactly where mitigation buys the most confidence.

### Cell six: a structured status table, and AI

The final cell writes a status table: CPI, SPI, cost variance, EAC and TCPI by account, plus the P50 and P80 finish. This becomes the single reviewed source for the monthly narrative. If your organisation allows it, you can give this table, and only this table, to an approved enterprise AI assistant with a prompt that says: draft a short narrative, cite the column for every number, and don't state causes that aren't in the data. It will produce a decent first draft in seconds. Then the controls lead checks every figure against the notebook, confirms causes with the control account managers and records what changed. The AI saves time on drafting. The person owns the forecast.

### Common mistakes and extending the lab

A few mistakes to avoid. Averaging CPIs across accounts: always compute from summed earned value and actual cost, so big accounts carry their proper weight. Running Monte Carlo on a network with missing logic, which produces precise-looking nonsense. Reporting P80 to two decimal places of a day, which implies precision you don't have. And sending an AI-drafted narrative without checking every number. When you're ready to extend the lab, replace the typed data with your real exports, add lags and start-to-start and finish-to-finish links, map more risks from your register, and link time-dependent costs to duration so schedule risk flows into cost risk. And validate: your notebook's deterministic finish should match your scheduling tool's for the same network.

### Recap and try this now

Let's recap. You've built a notebook that loads earned value data by account and period, calculates the metrics correctly, produces an EAC range and TCPI, flags a worrying trend with a simple rule, runs a Monte Carlo schedule risk analysis with PERT ranges, correlation and a discrete risk, identifies the critical drivers, and writes a structured status table for human review. Here's your try-this-now. Build the notebook from the lesson text and run it. Then change one thing: widen the maximum duration for procurement from seventy to ninety days, and run it again. How far does P80 move? And does the plan confidence change? That single experiment will teach you more about schedule risk than any histogram in a report.

## Key takeaways

- Keep EVM data as one row per control account per period, with cumulative PV, EV, AC and BAC.
- Compute project indices from summed EV, PV and AC; never average account CPIs.
- PERT ranges, a shared productivity factor and discrete register risks make a Monte Carlo model realistic.
- Read plan confidence, P50, P80, criticality and a tornado ranking together to target mitigation.
- A structured, reviewed status table is the only input an approved AI tool should draft narratives from.

## Try it

Build the notebook from the lesson, run it, then widen the procurement maximum from 70 to 90 days. Record how P80 and plan confidence change and write a two-sentence explanation.

- [Previous: Quantitative risk analysis and Monte Carlo simulation](https://optimizeall.com/learn/project-controls-with-ai/monte-carlo-and-quantitative-risk)
- [Next: Change control and baseline management](https://optimizeall.com/learn/project-controls-with-ai/change-control)
- [All lessons of Project Controls in the AI Era](https://optimizeall.com/learn/project-controls-with-ai)
