Advanced Prompt EngineeringPrompt operations: versioning, cost and latency · Lesson 15 of 17

Versioning and managing prompts in production

Article · 11 min · 7 min lecture

Video lecture

Versioning and managing prompts in production

11 chapters · about 7 min · full transcript

Coming soon

Chapter 1 of 11

Versioning prompts in production

  • Prompts are code
  • What to version
  • The change workflow
  • Model migrations and observability

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Prompts are code

In a production AI feature, the prompt is as much a part of the system as any function. A one-word edit can change behaviour for thousands of users. Yet prompts often live in chat histories, shared documents or someone's head. Prompt operations applies software discipline to prompts: version control, review, testing, staged release and rollback.

What to version

A "prompt" in production is usually a bundle:

  • system prompt and templates (with named variables)
  • few-shot examples
  • output schema
  • model identifier and parameters (temperature, maximum tokens, reasoning effort, tools enabled)
  • retrieval settings, if any (which index, how many passages)

Change any one and outputs can change, so version the bundle together:

id: lead-reply-drafter
version: 2.3.0
model: <provider-model-id>          # pin the exact model identifier
params: { temperature: 0.3, max_output_tokens: 400 }
system: prompts/lead-reply/system.v2.3.md
examples: prompts/lead-reply/examples.v2.json
schema: schemas/lead-reply.v1.json
changelog:
  - 2.3.0: Added Arabic formality rule after 4 complaints (eval +2 on 'tone' tag)
  - 2.2.1: Removed 'always ask about budget' - caused pushy tone

A simple semantic versioning convention

  • Major (3.0.0): changes the output contract, such as schema fields, that downstream code relies on.
  • Minor (2.4.0): behaviour changes: new rules, new examples, new model.
  • Patch (2.3.1): wording fixes that should not change behaviour (verify with evals).

The change workflow

  1. Propose a change with a reason, linked to failing cases or user feedback.
  2. Run the evaluation set on the old and new versions. Compare by tag, not just overall.
  3. Review like code: a second person reads the diff and the eval results.
  4. Stage the release: shadow mode (run both, show only old), or a small percentage of traffic, before everyone.
  5. Monitor live metrics and be ready to roll back instantly by switching the version pointer.

Model upgrades are prompt changes

Providers regularly release new models and retire old ones. A new model can follow instructions more literally, format differently, reason more or less by default, or react differently to emphasis. Treat an upgrade exactly like a prompt change: run evals, expect to adjust instructions, and check for over- or under-correction of old workaround rules. Pin exact model identifiers in production, rather than floating aliases, so behaviour does not change underneath you without an eval run. Keep track of announced deprecation dates for the models you use.

Separating prompt content from code

Store prompts as files or in a prompt management system rather than as long strings inside application code. Benefits: non-engineers (product, content, support leads) can propose improvements; diffs are readable; the same prompt can be tested in isolation. Whatever tool you use, the key properties are: history, diffs, environment separation (development, staging, production) and the ability to pin a version.

Worked example

A support team's reply prompt was edited directly in production by three people over a month. Satisfaction dipped, and nobody knew why. After moving to versioned prompts with a changelog and eval runs, they bisected the history: a patch that added "keep replies under 60 words" had truncated refund explanations. Rolling back that single change recovered quality. Without versioning, the regression would have been a mystery.

Observability

Log for every request: prompt version, model, parameters, input size, output size, latency and any grader or user feedback signals. This lets you answer questions like "did version 2.3 increase escalations?" instead of guessing. Handle logged content according to your privacy policy; mask personal data where you can.

Hands-on: a prompt registry in your repository

A lightweight pattern that works for most teams: prompts live in the repo, a manifest pins everything, and a CI job runs the eval set on every pull request that touches a prompt.

# prompt_registry.py
import hashlib, json, pathlib, yaml  # pip install pyyaml

ROOT = pathlib.Path("prompts")

def load(prompt_id: str, version: str) -> dict:
    manifest = yaml.safe_load((ROOT / prompt_id / f"{version}.yaml").read_text(encoding="utf-8"))
    system = (ROOT / prompt_id / manifest["system"]).read_text(encoding="utf-8")
    manifest["system_text"] = system
    manifest["fingerprint"] = hashlib.sha256(
        json.dumps({"system": system, "model": manifest["model"], "params": manifest["params"]},
                   sort_keys=True).encode()).hexdigest()[:12]
    return manifest

bundle = load("lead-reply-drafter", "2.3.0")
# log bundle["fingerprint"] with every request so incidents can be traced to an exact bundle
# .github/workflows/prompt-evals.yml (sketch)
on:
  pull_request:
    paths: ["prompts/**", "evals/**"]
jobs:
  evals:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pip install anthropic pyyaml
      - run: python evals/run.py --compare main --fail-on-regression 2
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}

The fingerprint catches silent changes: if two requests claim the same version but have different fingerprints, someone edited a file without bumping the version.

Model migrations deserve their own checklist

Providers retire models on published schedules, and the newest generations can change request shapes as well as behaviour (for example new effort controls, removed parameters or deprecated features such as prefill). For each migration:

  1. Read the provider's migration guide and deprecations page.
  2. Update request code for removed or renamed parameters.
  3. Re-run the full eval set; compare by tag.
  4. Audit the prompt for legacy workarounds that now over-correct.
  5. Re-baseline token counts and costs, since tokenisers can change between generations.
  6. Stage the rollout and keep the old version pinned for rollback until the retirement date.

Going further

Add continuous evaluation: a scheduled job runs your eval set against the production version daily or weekly, catching silent drift (for example from provider-side changes or data changes in retrieval). Combine this with sampled human review of live outputs for dimensions your graders cannot measure.

Key takeaways

  • Version the whole bundle: prompt, examples, schema, model ID, parameters and retrieval settings.
  • Use a change workflow: reason, evals on old vs new, review, staged release, monitoring and rollback.
  • Treat model upgrades as prompt changes; pin exact model IDs and track deprecations.
  • Log prompt version and key metrics per request so regressions can be traced.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why pin an exact model identifier in production instead of a 'latest' alias?
  2. Changing the output schema's field names in a way downstream code depends on is which kind of version change?
  3. What is shadow mode?

Put it into practice

Create a version manifest (like the YAML above) for one prompt you use, with a changelog entry explaining why each rule exists.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.