Building AI Products & WorkflowsPrototyping and build vs buy · Lesson 3 of 18

Rapid prototyping of AI features

Article · 11 min · 9 min lecture

Video lecture

Rapid prototyping of AI features

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Rapid prototyping

  • Why paper plans fail for AI
  • The four-rung ladder
  • The spreadsheet evaluation

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why prototype first

AI features are unusually hard to judge on paper. A prompt that seems obviously right may fail on real inputs; a use case that sounds marginal may turn out delightful. Prototypes answer the key questions cheaply: Can the model do this well enough, on our data, in a way users find valuable?

The prototyping ladder

  1. Chat prototype (hours). Try the task manually in an AI assistant with 10 to 20 real examples. Refine the prompt. This alone kills or confirms many ideas.
  2. Prompt + spreadsheet evaluation (days). Run a fixed prompt over 50 to 100 real inputs (many tools and simple scripts can do this), and have experts grade the outputs in a spreadsheet.
  3. Clickable or no-code prototype (days to weeks). A simple interface, workflow tool or internal app so real users can try it in context.
  4. Pilot (weeks). A limited production version with a small group of users, logging, and evaluation.

Move up only when the lower rung shows promise. Each rung costs roughly an order of magnitude more than the one below.

What to learn at each stage

  • Feasibility: does output quality meet the "good enough" bar on real, messy inputs?
  • Value: do users save time or produce better work? Do they want to use it?
  • Workflow fit: where does it slot into existing tools and steps?
  • Risks: what failure modes appear? How often? Are they detectable?
  • Cost signals: tokens per task, latency, review time.

Using real data safely

Prototypes are only informative with realistic inputs, but prototyping often happens before security reviews. Balance this by:

  • Using approved tools for any sensitive data.
  • Anonymising or synthesising data where possible (while checking that synthetic data isn't unrealistically clean).
  • Starting with public or low-sensitivity examples.

Prompt and evaluation assets are the real output

The most valuable outputs of prototyping are not the demo, but:

  • a tested prompt (or chain) with versions
  • an evaluation set of real inputs with expert grades
  • a documented list of failure modes
  • user feedback and observed workflow fit

These carry directly into the production build (or vendor evaluation).

Worked example: a proposal assistant

An agency tests whether AI can draft proposal sections.

  • Chat prototype: using five past briefs, the team finds drafts are generic unless given case studies and service descriptions. They build a context pack.
  • Spreadsheet evaluation: 30 past briefs run through the prompt with context; senior staff score drafts 1 to 5 on relevance and accuracy. Most scores are acceptable for first drafts; case-study selection is a common failure.
  • No-code prototype: a form where account managers paste the brief and get a draft in a shared document. Usage over three weeks is high; edit time is measured.
  • Decision: proceed to a pilot with retrieval over the case-study library to fix selection errors.

Total investment before any engineering: modest, with strong evidence for the next step.

Common prototyping mistakes

  • Demoing on hand-picked easy examples.
  • Skipping users and testing only with the builders.
  • Treating a prototype as production (no logging, no security, no evaluation).
  • Falling in love with a demo and ignoring the evaluation data.

Who should build prototypes?

Early rungs rarely need engineers. Product managers, analysts and domain experts can run chat prototypes and spreadsheet evaluations themselves, which keeps iteration fast and puts domain judgement at the centre. Bring in engineering when you reach integration, security and scale questions, with the evaluation evidence already in hand.

Hands-on: the spreadsheet evaluation runner

Rung 2 of the ladder ("prompt + spreadsheet evaluation") is where most ideas are proven or killed. This script runs one prompt over every row of a CSV of real inputs and writes the outputs to a new CSV with empty grading columns for your experts. It uses the Claude API; the pattern is identical with other providers.

pip install anthropic
export ANTHROPIC_API_KEY=...          # use an approved account for real data
import csv, os, time
import anthropic

client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")   # check the current model list
PROMPT_VERSION = "proposal-draft-v3"
SYSTEM = open("prompts/proposal-draft-v3.md", encoding="utf-8").read()   # keep prompts in files, versioned

def run_one(text: str) -> str:
    for attempt in range(4):
        try:
            resp = client.messages.create(model=MODEL, max_tokens=3000, system=SYSTEM,
                                          messages=[{"role": "user", "content": text}])
            return "".join(b.text for b in resp.content if b.type == "text")
        except anthropic.RateLimitError:
            time.sleep(2 ** attempt * 5)             # back off and retry
        except anthropic.APIStatusError as e:
            return f"ERROR {e.status_code}"
    return "ERROR rate_limited"

with open("inputs.csv", encoding="utf-8") as f_in, open(f"outputs-{PROMPT_VERSION}.csv", "w", newline="", encoding="utf-8") as f_out:
    reader = csv.DictReader(f_in)
    writer = csv.DictWriter(f_out, fieldnames=["id", "input", "output", "prompt_version",
                                               "grade_relevance_1to5", "grade_accuracy_1to5", "failure_mode", "grader"])
    writer.writeheader()
    for row in reader:
        writer.writerow({"id": row["id"], "input": row["input"], "output": run_one(row["input"]),
                         "prompt_version": PROMPT_VERSION})

For hundreds of rows where you do not need answers immediately, most providers offer a batch API at a discount (Anthropic's Message Batches API, for example, processes requests asynchronously at lower cost); submit the rows as a batch and collect results when it ends.

Grading tips: give graders a written rubric with an example for each score; have two graders score the first 20 rows independently and discuss disagreements; record a failure_mode label in a few words ("wrong case study", "invented figure", "tone too formal"). Those labels become your prompt-improvement backlog and, later, your evaluation categories.

Rung 3 today: AI-assisted prototypes

Rung 3 got much cheaper with AI coding agents and app builders, which can produce a clickable internal tool from a written spec in hours. The lesson "Prototyping with AI coding tools" covers how to use them without mistaking a prototype for production.

Going further

Timebox each rung and define the evidence needed to move up before you start ("at least 70% of drafts rated 4+ by two senior reviewers" or "users choose the tool for at least half of eligible tasks"). This keeps prototyping honest and fast.

Key takeaways

  • Prototype to test feasibility, value, workflow fit, risk and cost on real inputs.
  • Climb the ladder: chat prototype, spreadsheet evaluation, no-code prototype, pilot, each only when justified.
  • Use real data safely: approved tools, anonymisation, low-sensitivity examples first.
  • The durable outputs are the tested prompt, evaluation set, failure modes and user evidence.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What is the cheapest first test of an AI idea?
  2. Why is demoing on hand-picked examples risky?
  3. Which prototype output is most valuable for the production build?

Put it into practice

Run a chat prototype for one use case on 15 real examples. Record failure modes and write the evidence threshold required to move to the next rung.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.