Skip to content

Advanced Prompt Engineering · Evaluation sets and grading · lesson 14 of 17 · 13 min

Automated prompt optimisation and auto-prompting

Why automate prompt improvement?

Manual prompt iteration works, but it is slow and biased toward the cases you happened to look at. Automated prompt optimisation (sometimes called auto-prompting) uses a model to propose prompt changes, scores each candidate on an evaluation set, and keeps the improvements. It turns prompt engineering into a search problem guided by your metric.

It is only as good as two things: your evaluation set and your metric. With a weak metric, an optimiser will happily find prompts that game it.

The family of techniques

| Approach | What it does | Good for | |---|---|---| | Prompt improvers in vendor consoles | Rewrite a draft prompt using best practices (structure, examples, clarity) | Fast first drafts; porting prompts between model families | | Meta-prompting loop (DIY) | A strong model reads failures and proposes an edited prompt; you score it | Small teams; full control; any provider | | Framework optimisers | Libraries such as DSPy compile programs and optimise instructions and examples against a metric; reflective optimisers such as GEPA analyse failures in natural language to propose better instructions | Multi-step pipelines; systematic search | | Example selection | Automatically choose which few-shot examples to include | Classification and extraction tasks |

The Anthropic Console includes a prompt improver and prompt generator, and the OpenAI platform includes a prompt optimizer in its playground; both are useful starting points, and both should be followed by evaluation on your own data.

Hands-on: a meta-prompting optimisation loop

This loop uses a strong model as the "optimiser" and your eval set as the judge. It keeps a train split for proposing changes and a held-out split for the final decision.

import json, os, random, anthropic
client = anthropic.Anthropic()
TASK_MODEL = os.environ.get("TASK_MODEL", "claude-haiku-4-5")   # the model you will deploy
OPT_MODEL = os.environ.get("OPT_MODEL", "claude-opus-5")         # the model proposing edits

def text(resp):
    return "".join(b.text for b in resp.content if b.type == "text").strip()

def classify(prompt, message):
    r = client.messages.create(model=TASK_MODEL, max_tokens=20, system=prompt,
                               messages=[{"role": "user", "content": message}])
    return text(r).upper()

def score(prompt, cases):
    results = [(c, classify(prompt, c["input"])) for c in cases]
    acc = sum(out == c["label"] for c, out in results) / len(cases)
    failures = [{"input": c["input"], "expected": c["label"], "got": out}
                for c, out in results if out != c["label"]]
    return acc, failures

def propose(prompt, failures):
    r = client.messages.create(model=OPT_MODEL, max_tokens=4000, messages=[{"role": "user", "content":
        f"""You improve classification prompts. Current prompt:
<prompt>{prompt}</prompt>
Failures on training cases:
<failures>{json.dumps(failures[:15], ensure_ascii=False)}</failures>
Diagnose the pattern behind the failures, then return an improved prompt inside
<new_prompt> tags. Keep the label set and output format unchanged. Do not copy
training inputs into the prompt."""}])
    out = text(r)
    return out.split("<new_prompt>")[1].split("</new_prompt>")[0].strip()

cases = [json.loads(l) for l in open("evals/leads.jsonl", encoding="utf-8")]
random.seed(7); random.shuffle(cases)
train, holdout = cases[: len(cases) * 2 // 3], cases[len(cases) * 2 // 3 :]

best = open("prompts/lead_classifier.md", encoding="utf-8").read()
best_acc, failures = score(best, train)
for round_ in range(5):
    candidate = propose(best, failures)
    acc, cand_failures = score(candidate, train)
    print(f"round {round_}: train acc {acc:.0%} (best {best_acc:.0%})")
    if acc > best_acc:
        best, best_acc, failures = candidate, acc, cand_failures

base_hold, _ = score(open("prompts/lead_classifier.md", encoding="utf-8").read(), holdout)
new_hold, _ = score(best, holdout)
print(f"holdout: baseline {base_hold:.0%} -> optimised {new_hold:.0%}")

Only adopt the new prompt if the held-out score improves by more than run-to-run noise (run each score several times if outputs vary), and read the new prompt yourself: optimisers sometimes insert odd rules or overfit to training quirks.

Framework optimisers in brief

DSPy expresses a pipeline as modules with typed signatures (for example question -> answer) and lets an optimiser tune the instructions and demonstrations against your metric; its GEPA optimiser uses a reflection model to read failure traces and propose improved instructions. A sketch (check the current DSPy documentation for exact signatures and model strings):

import dspy
dspy.configure(lm=dspy.LM("anthropic/claude-haiku-4-5"))
classify = dspy.Predict("message -> label")

def metric(example, pred, trace=None, pred_name=None, pred_trace=None):
    return float(pred.label.strip().upper() == example.label)

optimizer = dspy.GEPA(metric=metric, auto="light",
                      reflection_lm=dspy.LM("anthropic/claude-opus-5"))
optimized = optimizer.compile(classify, trainset=trainset, valset=valset)

Framework optimisers shine on multi-step pipelines where hand-tuning each step is impractical. The trade-off is another dependency and less direct control over the final prompt text.

A cost-saving use: make small models punch above their weight

A common, valuable pattern is optimising the prompt for a smaller, cheaper deployment model using a stronger model as optimiser or reflector. If the optimised small model meets your accuracy bar, you keep the savings on every future request.

Guard rails for optimisation

  • Never optimise on your test set. Use train and validation splits; touch the held-out test set only for the final decision.
  • Constrain what can change. Lock the output format, label set and safety rules; let the optimiser edit instructions and examples only.
  • Watch for metric gaming. If your metric only checks format, you will get perfectly formatted wrong answers.
  • Review before shipping. Treat an optimised prompt like any other prompt change: diff, review, eval by tag, staged rollout.

How to know it worked

The optimised prompt beats the baseline on held-out data by a margin larger than run-to-run variation, does not regress any important tag, and a human reviewer can explain why its changes make sense.

Video lecture: Automated prompt optimisation and auto-prompting

Lecture coming soon · 11 chapters · about 7 minutes. Read the full transcript below.

  1. Automated prompt optimisation
  2. The idea and the warning
  3. The technique family
  4. The meta-prompting loop
  5. Decide on held-out data
  6. Framework optimisers
  7. Make small models punch above their weight
  8. Example 1: email routing
  9. Example 2: small model, big savings (illustrative)
  10. Guard rails
  11. Recap

Lecture transcript

Automated prompt optimisation

What if you could hand your prompt, your eval set and a metric to a model, and come back to a better prompt? That is the promise of automated prompt optimisation, sometimes called auto-prompting. It works, with one big condition: it is only as good as your evaluation set and your metric. In this lecture you will learn the main techniques, build a meta-prompting loop with proper train and held-out splits, see how framework optimisers like DSPy and GEPA fit in, and learn the guard rails that stop an optimiser from gaming your metric.

The idea and the warning

Manual iteration works, but it is slow and biased toward the cases you happened to look at. Automated optimisation uses a model to propose prompt changes, scores each candidate on an evaluation set, and keeps the improvements. It turns prompt engineering into a search problem guided by your metric. And here is the warning in one sentence: with a weak metric, an optimiser will happily find prompts that game it. If your metric only checks that the JSON is valid, you will get beautifully formatted wrong answers.

The technique family

There is a family of techniques. Vendor console prompt improvers rewrite a draft using best practices; the Anthropic Console has a prompt improver and generator, and the OpenAI platform has a prompt optimizer in its playground. They are great for first drafts and for porting prompts between model families. A do-it-yourself meta-prompting loop has a strong model read failures and propose edits, which you score. Framework optimisers such as DSPy tune instructions and examples against a metric. And example-selection methods automatically choose which few-shot examples to include.

The meta-prompting loop

Let us build the do-it-yourself loop from the lesson. You have a labelled dataset. Shuffle it and split it: two thirds for training, one third held out. The task model is the one you will deploy, perhaps a small, fast one. The optimiser model is a stronger one. Score the current prompt on the training cases and collect failures. Send the optimiser the current prompt and up to fifteen failures, and ask it to diagnose the pattern and return an improved prompt, keeping the labels and format unchanged and not copying training inputs. Score the candidate. Keep it only if it beats the best so far. Repeat for a few rounds.

Decide on held-out data

Then the decision that matters. Score the original and the optimised prompt on the held-out set, which the optimiser never saw. Adopt the new prompt only if it improves by more than run-to-run noise, so run each score more than once if outputs vary. And read the new prompt yourself. Optimisers sometimes insert odd rules or overfit to quirks of the training data. Treat an optimised prompt like any other prompt change: diff, review, evaluate by tag, and stage the rollout.

Framework optimisers

Framework optimisers take this further. DSPy expresses a pipeline as modules with typed signatures, like message to label, and lets an optimiser tune instructions and demonstrations against your metric. Its GEPA optimiser uses a reflection model to read failure traces in natural language and propose better instructions. The lesson includes a short sketch; check the current DSPy documentation for exact signatures. Frameworks shine on multi-step pipelines where hand-tuning every step is impractical. The trade-off is another dependency and less direct control over the final prompt text.

Make small models punch above their weight

One of the most valuable uses is making small models punch above their weight. Use a strong model as the optimiser or reflector, and optimise the prompt for a smaller, cheaper deployment model. If the optimised small model meets your accuracy bar, you keep the savings on every future request. For a high-volume classifier, that can change the economics of the whole feature.

Example 1: email routing

A simple worked example. You have a prompt that classifies customer emails into five departments, and thirty labelled examples. You run the meta-prompting loop. The optimiser reads the failures and notices a pattern: emails about invoices for damaged goods keep going to billing when they belong to returns. It proposes a new instruction: if an email mentions damage or a faulty item, route to returns even if it mentions an invoice. Training accuracy improves. More importantly, accuracy on the ten held-out emails also improves. You read the new rule, it makes sense, and you adopt it. The optimiser found a rule you might have missed, and you verified it before shipping.

Example 2: small model, big savings (illustrative)

Now a business scenario, with illustrative numbers. A marketplace in Karachi tags about fifty thousand product listings a month into two hundred categories. They use a large model today, which is accurate but expensive. They want to move to a smaller model. With the small model and the current prompt, held-out accuracy is around eighty-one percent, against ninety-two for the large model. They run an optimisation loop for six rounds, using the large model as the optimiser, a training set of six hundred labelled listings and a held-out set of three hundred. The optimised small-model prompt reaches about ninety percent on held-out data. The team locks the category list and output format, reviews the final prompt, and routes low-confidence items to the large model. Monthly cost drops by well over half. Numbers illustrative, but this is the most valuable use of optimisation in practice.

Guard rails

Four guard rails. Never optimise on your test set; use train and validation data, and touch the held-out test only for the final decision. Constrain what can change: lock the output format, label set and safety rules, and let the optimiser edit instructions and examples only. Watch for metric gaming, and make your metric measure what you really care about. And review before shipping. An optimised prompt that nobody can explain is a liability.

Recap

To recap. Automated optimisation proposes changes and keeps those that improve a metric on an eval set. Options range from console improvers to meta-prompting loops and framework optimisers like DSPy with GEPA. Optimise on training data, decide on held-out data, lock the contract and safety rules, and review. Try this now: split a labelled dataset, run the loop from the lesson for three rounds, and report baseline versus optimised held-out accuracy with a short note on what changed and why. Next module: prompt operations, starting with versioning.

Key takeaways

  • Automated optimisation proposes prompt changes and keeps those that improve a metric on an eval set.
  • Options range from console prompt improvers to DIY meta-prompting loops and framework optimisers like DSPy with GEPA.
  • Optimise on train data, decide on held-out data, and lock format and safety rules.
  • A strong optimiser model can tune prompts for a cheaper deployment model; always review and stage the result.

Try it

Split a labelled dataset into train and holdout, run the meta-prompting loop from the lesson for three rounds, and report baseline versus optimised holdout accuracy with a short review of what changed.