Advanced Prompt EngineeringPrompt operations: versioning, cost and latency · Lesson 15 of 17
Versioning and managing prompts in production
Video lecture
Versioning and managing prompts in production
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Versioning prompts in production
Customer satisfaction dipped for a month and nobody knew why. Three people had been editing the support prompt directly in production. When the team finally moved to versioned prompts and bisected the history, they found it: a tiny patch that said keep replies under sixty words had been truncating every refund explanation. In this lecture you will learn to treat prompts as code: what to version, a simple versioning convention, the change workflow, how to handle model migrations, and the logging that makes regressions traceable.
0:37 Why version prompts?
Why does versioning matter? Because prompts change more often than code, and the changes are harder to see. A single word, like always, can alter thousands of outputs. Without history, every regression becomes a mystery, and nobody dares touch the prompt, so it slowly rots. With versioning, every change has a reason, a test result and an undo button. Think of it as the difference between editing a shared document with track changes, and editing one without. Here is the key idea for this lecture: a prompt in production is not a sentence, it is a component.
1:19 Version the bundle
In a production AI feature, the prompt is as much a part of the system as any function. A one-word edit can change behaviour for thousands of users. Yet prompts often live in chat histories, shared documents or someone's head. And a production prompt is really a bundle: the system prompt and templates, few-shot examples, the output schema, the model identifier and parameters like maximum tokens, reasoning effort and enabled tools, and retrieval settings. Change any one and outputs can change, so version the bundle together in a manifest.
1:58 Semantic versioning for prompts
A simple convention helps. Major version: changes the output contract, like schema fields downstream code relies on. Minor version: behaviour changes, such as new rules, new examples, or a new model. Patch: wording fixes that should not change behaviour, and you verify that with evals. Keep a changelog entry for every change: what changed, why, and what the evals showed. For example, added an Arabic formality rule after four complaints, tone tag improved by two cases.
2:31 The change workflow
Here is the change workflow. Propose a change with a reason, linked to failing cases or user feedback. Run the evaluation set on old and new versions, and compare by tag, not just overall. Review it like code: a second person reads the diff and the eval results. Stage the release, in shadow mode where both run but only the old one is shown, or to a small percentage of traffic. Then monitor live metrics, and be ready to roll back instantly by switching the version pointer.
3:09 Hands-on: prompt registry + CI
The lesson's hands-on registry keeps prompts in your repository with a YAML manifest per version, and computes a short fingerprint of the system text, model and parameters. You log that fingerprint with every request. If two requests claim the same version but have different fingerprints, someone edited a file without bumping the version. A continuous integration job runs your eval set on every pull request that touches a prompt, and fails the build if a tag regresses beyond a threshold.
3:44 Model migrations
Model upgrades are prompt changes. Providers release new models and retire old ones on published schedules, and the newest generations can change request shapes as well as behaviour: new effort controls, removed parameters, deprecated features such as prefill. So for each migration, read the provider's migration guide and deprecations page, update request code, re-run the full eval set by tag, audit the prompt for legacy workarounds that now over-correct, re-baseline token counts and costs because tokenisers can change, and stage the rollout with the old version pinned until the retirement date. Pin exact model identifiers in production rather than floating aliases.
4:28 Observability
Observability ties it together. For every request, log the prompt version or fingerprint, model, parameters, input and output sizes, latency, and any grader or user feedback signals. Then you can answer questions like did version two point three increase escalations, instead of guessing. Handle logged content according to your privacy policy and mask personal data where you can. And add continuous evaluation: a scheduled job runs your eval set against production daily or weekly to catch silent drift.
5:02 Example 1: a 'small' wording change
A simple worked example. Your summariser prompt is version one point two point zero. A colleague changes, be concise, to, maximum three bullets. Is that a patch? It sounds like wording, but it changes behaviour, so it is a minor version: one point three point zero. You run the eval set: overall score is unchanged, but the tag long meetings drops, because three bullets cannot hold all decisions. You write that in the changelog, adjust to, up to five bullets for meetings over an hour, and re-run. Now the history explains itself, and anyone can see why the rule exists.
5:45 Example 2: a model migration (illustrative)
Now a business scenario, with illustrative numbers. A fintech in London runs about twenty AI prompts in production across onboarding, support and fraud notes. After a provider announces a model retirement, they must migrate everything within three months. Because every prompt is a versioned bundle with a pinned model, an eval set and a changelog, they run each bundle against the new model in a day. Seventeen pass with small differences. Three regress on specific tags; one support prompt, for example, has a legacy rule, always restate the question, that now makes replies longer. They remove it, re-run, and pass. They also re-baseline token counts, because the new model's tokeniser differs, and update their cost forecasts. Staged rollouts complete in three weeks, with the old versions pinned for rollback. Illustrative figures, but notice how the preparation made the migration boring, which is exactly what you want.
6:48 Recap
To recap. A production prompt is a bundle: prompt, examples, schema, model and parameters, and retrieval settings. Version it together, with a changelog of reasons. Follow a workflow of evals, review, staged release and rollback, treat model upgrades as prompt changes, pin model IDs, and log versions with every request. Try this now: create a version manifest for one prompt you use, with a changelog entry explaining why each rule exists. Next: cost and latency trade-offs.
Prompts are code
In a production AI feature, the prompt is as much a part of the system as any function. A one-word edit can change behaviour for thousands of users. Yet prompts often live in chat histories, shared documents or someone's head. Prompt operations applies software discipline to prompts: version control, review, testing, staged release and rollback.
What to version
A "prompt" in production is usually a bundle:
- system prompt and templates (with named variables)
- few-shot examples
- output schema
- model identifier and parameters (temperature, maximum tokens, reasoning effort, tools enabled)
- retrieval settings, if any (which index, how many passages)
Change any one and outputs can change, so version the bundle together:
id: lead-reply-drafter
version: 2.3.0
model: <provider-model-id> # pin the exact model identifier
params: { temperature: 0.3, max_output_tokens: 400 }
system: prompts/lead-reply/system.v2.3.md
examples: prompts/lead-reply/examples.v2.json
schema: schemas/lead-reply.v1.json
changelog:
- 2.3.0: Added Arabic formality rule after 4 complaints (eval +2 on 'tone' tag)
- 2.2.1: Removed 'always ask about budget' - caused pushy toneA simple semantic versioning convention
- Major (3.0.0): changes the output contract, such as schema fields, that downstream code relies on.
- Minor (2.4.0): behaviour changes: new rules, new examples, new model.
- Patch (2.3.1): wording fixes that should not change behaviour (verify with evals).
The change workflow
- Propose a change with a reason, linked to failing cases or user feedback.
- Run the evaluation set on the old and new versions. Compare by tag, not just overall.
- Review like code: a second person reads the diff and the eval results.
- Stage the release: shadow mode (run both, show only old), or a small percentage of traffic, before everyone.
- Monitor live metrics and be ready to roll back instantly by switching the version pointer.
Model upgrades are prompt changes
Providers regularly release new models and retire old ones. A new model can follow instructions more literally, format differently, reason more or less by default, or react differently to emphasis. Treat an upgrade exactly like a prompt change: run evals, expect to adjust instructions, and check for over- or under-correction of old workaround rules. Pin exact model identifiers in production, rather than floating aliases, so behaviour does not change underneath you without an eval run. Keep track of announced deprecation dates for the models you use.
Separating prompt content from code
Store prompts as files or in a prompt management system rather than as long strings inside application code. Benefits: non-engineers (product, content, support leads) can propose improvements; diffs are readable; the same prompt can be tested in isolation. Whatever tool you use, the key properties are: history, diffs, environment separation (development, staging, production) and the ability to pin a version.
Worked example
A support team's reply prompt was edited directly in production by three people over a month. Satisfaction dipped, and nobody knew why. After moving to versioned prompts with a changelog and eval runs, they bisected the history: a patch that added "keep replies under 60 words" had truncated refund explanations. Rolling back that single change recovered quality. Without versioning, the regression would have been a mystery.
Observability
Log for every request: prompt version, model, parameters, input size, output size, latency and any grader or user feedback signals. This lets you answer questions like "did version 2.3 increase escalations?" instead of guessing. Handle logged content according to your privacy policy; mask personal data where you can.
Hands-on: a prompt registry in your repository
A lightweight pattern that works for most teams: prompts live in the repo, a manifest pins everything, and a CI job runs the eval set on every pull request that touches a prompt.
# prompt_registry.py
import hashlib, json, pathlib, yaml # pip install pyyaml
ROOT = pathlib.Path("prompts")
def load(prompt_id: str, version: str) -> dict:
manifest = yaml.safe_load((ROOT / prompt_id / f"{version}.yaml").read_text(encoding="utf-8"))
system = (ROOT / prompt_id / manifest["system"]).read_text(encoding="utf-8")
manifest["system_text"] = system
manifest["fingerprint"] = hashlib.sha256(
json.dumps({"system": system, "model": manifest["model"], "params": manifest["params"]},
sort_keys=True).encode()).hexdigest()[:12]
return manifest
bundle = load("lead-reply-drafter", "2.3.0")
# log bundle["fingerprint"] with every request so incidents can be traced to an exact bundle# .github/workflows/prompt-evals.yml (sketch)
on:
pull_request:
paths: ["prompts/**", "evals/**"]
jobs:
evals:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pip install anthropic pyyaml
- run: python evals/run.py --compare main --fail-on-regression 2
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}The fingerprint catches silent changes: if two requests claim the same version but have different fingerprints, someone edited a file without bumping the version.
Model migrations deserve their own checklist
Providers retire models on published schedules, and the newest generations can change request shapes as well as behaviour (for example new effort controls, removed parameters or deprecated features such as prefill). For each migration:
- Read the provider's migration guide and deprecations page.
- Update request code for removed or renamed parameters.
- Re-run the full eval set; compare by tag.
- Audit the prompt for legacy workarounds that now over-correct.
- Re-baseline token counts and costs, since tokenisers can change between generations.
- Stage the rollout and keep the old version pinned for rollback until the retirement date.
Going further
Add continuous evaluation: a scheduled job runs your eval set against the production version daily or weekly, catching silent drift (for example from provider-side changes or data changes in retrieval). Combine this with sampled human review of live outputs for dimensions your graders cannot measure.
Key takeaways
- Version the whole bundle: prompt, examples, schema, model ID, parameters and retrieval settings.
- Use a change workflow: reason, evals on old vs new, review, staged release, monitoring and rollback.
- Treat model upgrades as prompt changes; pin exact model IDs and track deprecations.
- Log prompt version and key metrics per request so regressions can be traced.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Create a version manifest (like the YAML above) for one prompt you use, with a changelog entry explaining why each rule exists.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.