Building AI Products & WorkflowsPrototyping and build vs buy · Lesson 3 of 18
Rapid prototyping of AI features
Video lecture
Rapid prototyping of AI features
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Rapid prototyping
AI features are strangely hard to judge on paper. A prompt that looks obviously right falls apart on real, messy inputs. A use case that sounds marginal turns out to be the thing everyone uses. So the cheapest way to make good AI product decisions is to prototype, fast, on real data. In this lesson you'll climb a four-rung prototyping ladder, learn what each rung should teach you, and run a spreadsheet evaluation that can kill or confirm an idea in days.
0:36 Analogy: taste before you print the menu
Why prototype before planning in detail? Here's an analogy. Chefs don't write a full menu and print it before tasting the dishes. They cook small batches, taste, adjust and only then commit. AI features are the same. A prompt that sounds perfect often tastes wrong on real inputs. Prototyping is tasting early, while changes are cheap.
1:00 The prototyping ladder
Here's the ladder. Rung one, a chat prototype, takes hours: try the task by hand in an AI assistant with ten to twenty real examples and refine the prompt. This alone kills or confirms many ideas. Rung two, prompt plus spreadsheet evaluation, takes days: run a fixed prompt over fifty to a hundred real inputs and have experts grade the outputs. Rung three, a clickable or no-code prototype, puts it in front of real users in context. Rung four is a pilot: a limited production version with logging and evaluation. Each rung costs roughly ten times the one below, so climb only when the lower rung shows promise.
1:47 What each rung must answer
At every rung you're answering five questions. Feasibility: is output quality good enough on real, messy inputs? Value: do users save time or produce better work, and do they want it? Workflow fit: where does it slot into existing tools? Risks: what failure modes appear, how often, and can you detect them? And cost signals: tokens per task, latency and review time. Write the evidence you need to move up before you start, for example at least seventy percent of drafts rated four or better by two senior reviewers.
2:26 Real data, safely
Prototypes only tell the truth with realistic inputs, but prototyping often happens before security reviews. Balance that. Use approved tools for any sensitive data. Anonymise or synthesise data where you can, while checking synthetic data isn't unrealistically clean. And start with public or low-sensitivity examples. Remember too that the most valuable outputs of prototyping aren't the demo. They're a tested, versioned prompt, an evaluation set of real inputs with expert grades, a documented list of failure modes, and real user feedback. All of that carries straight into the production build or the vendor evaluation.
3:07 Simple example: interview notes
A simple example of rung one. A recruitment agency wonders if AI can turn messy interview notes into a clean candidate summary. A recruiter pastes five real, anonymised sets of notes into an approved assistant with a draft prompt. Two summaries are great, two miss salary expectations, one invents a qualification. In twenty minutes they learn the prompt needs a field list and a rule about never adding facts. That's rung one doing its job.
3:40 Worked example: proposal assistant
Here's a worked example. An agency tests whether AI can draft proposal sections. In the chat prototype, five past briefs show drafts are generic unless the model gets case studies and service descriptions, so they build a context pack. In the spreadsheet evaluation, thirty past briefs run through the prompt and senior staff score relevance and accuracy. Most drafts are acceptable first drafts, but case-study selection often fails. A simple form prototype then gets heavy use for three weeks while they measure edit time. Decision: pilot, with retrieval over the case-study library to fix selection. Total spend before any engineering: modest.
4:24 Business example (illustrative)
Illustrative numbers for the proposal assistant. Thirty past briefs were scored by two senior staff: about seventy percent rated four or five for relevance, but case-study selection was wrong in roughly a third. In the three-week form prototype, account managers used it for most proposals and median drafting time for the approach section fell from about ninety minutes to about forty. That evidence, not the demo, got the pilot funded.
4:54 Common mistakes
Avoid four mistakes. Demoing on hand-picked easy examples. Testing only with the builders instead of real users. Treating a prototype as production, with no logging, security or evaluation. And falling in love with the demo while ignoring the evaluation data. Early rungs rarely need engineers. Product managers, analysts and domain experts can run chat prototypes and spreadsheet evaluations themselves, which keeps iteration fast and puts domain judgement at the centre.
5:24 Hands-on in the lesson
The hands-on section gives you a spreadsheet evaluation runner. It reads real inputs from a CSV, runs a versioned prompt through the Claude API with polite back-off on rate limits, and writes outputs to a new CSV with empty columns for relevance, accuracy, failure mode and grader. You'll also see when to use a batch API for big runs at lower cost, and grading tips: a written rubric, two graders on the first twenty rows, and short failure-mode labels that become your improvement backlog.
6:01 How you'll know it succeeded
How will you know a prototype phase succeeded? Not by how impressive the demo looked. By whether you hit the evidence bar you set in advance, whether you now have a documented list of failure modes, and whether you have a clear decision: climb to the next rung, change approach, or stop. A prototype that ends with let's see is a prototype that didn't finish.
6:29 Keep the loop fast
One more practical point: speed of feedback. The best prototype teams shorten the loop between trying something and seeing whether it worked. That means a small, fixed set of real inputs you can re-run in minutes, graders who are available that week, and a place where results live side by side across prompt versions. If each iteration takes days, you'll stop iterating.
6:56 Watch me do it: evaluation runner
Watch me do it. I open the evaluation runner. First, the system prompt loads from a versioned file, proposal draft v three, so every output records which prompt made it. Next, run one calls the API; on a rate-limit error it waits five, ten, twenty, then forty seconds before giving up, and on any other API error it records the status code instead of crashing. Then the main block reads inputs dot csv and writes outputs with columns for relevance, accuracy, failure mode and grader, left empty. I run it on thirty briefs. It takes a few minutes and produces outputs v three dot csv. I open it in a spreadsheet, share it with two senior staff, and ask each to grade the first twenty rows independently. We compare, agree on the rubric where we disagreed, and they grade the rest. Finally, I sort by failure mode, and wrong case study is the biggest group.
8:03 Recap
To recap: prototype to test feasibility, value, fit, risk and cost on real inputs. Climb the ladder one rung at a time with an evidence bar set in advance. Use real data safely, and keep the durable assets: prompt, evaluation set, failure modes and user evidence. Your next step is to run a chat prototype on fifteen real examples, record the failure modes, and write the threshold for moving to rung two. Next, we'll see how AI coding tools make rung three faster than ever.
8:40 Try this now (under an hour)
Try this now. Pick your scoped use case and collect fifteen real, safe examples. Run them through an approved AI assistant with your first prompt. For each, write one line: good, fixable, or failed, and why. Group the failures into two or three types, then write the evidence bar you'd need to see before moving to rung two. That's your prototype plan, done in under an hour.
Why prototype first
AI features are unusually hard to judge on paper. A prompt that seems obviously right may fail on real inputs; a use case that sounds marginal may turn out delightful. Prototypes answer the key questions cheaply: Can the model do this well enough, on our data, in a way users find valuable?
The prototyping ladder
- Chat prototype (hours). Try the task manually in an AI assistant with 10 to 20 real examples. Refine the prompt. This alone kills or confirms many ideas.
- Prompt + spreadsheet evaluation (days). Run a fixed prompt over 50 to 100 real inputs (many tools and simple scripts can do this), and have experts grade the outputs in a spreadsheet.
- Clickable or no-code prototype (days to weeks). A simple interface, workflow tool or internal app so real users can try it in context.
- Pilot (weeks). A limited production version with a small group of users, logging, and evaluation.
Move up only when the lower rung shows promise. Each rung costs roughly an order of magnitude more than the one below.
What to learn at each stage
- Feasibility: does output quality meet the "good enough" bar on real, messy inputs?
- Value: do users save time or produce better work? Do they want to use it?
- Workflow fit: where does it slot into existing tools and steps?
- Risks: what failure modes appear? How often? Are they detectable?
- Cost signals: tokens per task, latency, review time.
Using real data safely
Prototypes are only informative with realistic inputs, but prototyping often happens before security reviews. Balance this by:
- Using approved tools for any sensitive data.
- Anonymising or synthesising data where possible (while checking that synthetic data isn't unrealistically clean).
- Starting with public or low-sensitivity examples.
Prompt and evaluation assets are the real output
The most valuable outputs of prototyping are not the demo, but:
- a tested prompt (or chain) with versions
- an evaluation set of real inputs with expert grades
- a documented list of failure modes
- user feedback and observed workflow fit
These carry directly into the production build (or vendor evaluation).
Worked example: a proposal assistant
An agency tests whether AI can draft proposal sections.
- Chat prototype: using five past briefs, the team finds drafts are generic unless given case studies and service descriptions. They build a context pack.
- Spreadsheet evaluation: 30 past briefs run through the prompt with context; senior staff score drafts 1 to 5 on relevance and accuracy. Most scores are acceptable for first drafts; case-study selection is a common failure.
- No-code prototype: a form where account managers paste the brief and get a draft in a shared document. Usage over three weeks is high; edit time is measured.
- Decision: proceed to a pilot with retrieval over the case-study library to fix selection errors.
Total investment before any engineering: modest, with strong evidence for the next step.
Common prototyping mistakes
- Demoing on hand-picked easy examples.
- Skipping users and testing only with the builders.
- Treating a prototype as production (no logging, no security, no evaluation).
- Falling in love with a demo and ignoring the evaluation data.
Who should build prototypes?
Early rungs rarely need engineers. Product managers, analysts and domain experts can run chat prototypes and spreadsheet evaluations themselves, which keeps iteration fast and puts domain judgement at the centre. Bring in engineering when you reach integration, security and scale questions, with the evaluation evidence already in hand.
Hands-on: the spreadsheet evaluation runner
Rung 2 of the ladder ("prompt + spreadsheet evaluation") is where most ideas are proven or killed. This script runs one prompt over every row of a CSV of real inputs and writes the outputs to a new CSV with empty grading columns for your experts. It uses the Claude API; the pattern is identical with other providers.
pip install anthropic
export ANTHROPIC_API_KEY=... # use an approved account for real dataimport csv, os, time
import anthropic
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5") # check the current model list
PROMPT_VERSION = "proposal-draft-v3"
SYSTEM = open("prompts/proposal-draft-v3.md", encoding="utf-8").read() # keep prompts in files, versioned
def run_one(text: str) -> str:
for attempt in range(4):
try:
resp = client.messages.create(model=MODEL, max_tokens=3000, system=SYSTEM,
messages=[{"role": "user", "content": text}])
return "".join(b.text for b in resp.content if b.type == "text")
except anthropic.RateLimitError:
time.sleep(2 ** attempt * 5) # back off and retry
except anthropic.APIStatusError as e:
return f"ERROR {e.status_code}"
return "ERROR rate_limited"
with open("inputs.csv", encoding="utf-8") as f_in, open(f"outputs-{PROMPT_VERSION}.csv", "w", newline="", encoding="utf-8") as f_out:
reader = csv.DictReader(f_in)
writer = csv.DictWriter(f_out, fieldnames=["id", "input", "output", "prompt_version",
"grade_relevance_1to5", "grade_accuracy_1to5", "failure_mode", "grader"])
writer.writeheader()
for row in reader:
writer.writerow({"id": row["id"], "input": row["input"], "output": run_one(row["input"]),
"prompt_version": PROMPT_VERSION})For hundreds of rows where you do not need answers immediately, most providers offer a batch API at a discount (Anthropic's Message Batches API, for example, processes requests asynchronously at lower cost); submit the rows as a batch and collect results when it ends.
Grading tips: give graders a written rubric with an example for each score; have two graders score the first 20 rows independently and discuss disagreements; record a failure_mode label in a few words ("wrong case study", "invented figure", "tone too formal"). Those labels become your prompt-improvement backlog and, later, your evaluation categories.
Rung 3 today: AI-assisted prototypes
Rung 3 got much cheaper with AI coding agents and app builders, which can produce a clickable internal tool from a written spec in hours. The lesson "Prototyping with AI coding tools" covers how to use them without mistaking a prototype for production.
Going further
Timebox each rung and define the evidence needed to move up before you start ("at least 70% of drafts rated 4+ by two senior reviewers" or "users choose the tool for at least half of eligible tasks"). This keeps prototyping honest and fast.
Key takeaways
- Prototype to test feasibility, value, workflow fit, risk and cost on real inputs.
- Climb the ladder: chat prototype, spreadsheet evaluation, no-code prototype, pilot, each only when justified.
- Use real data safely: approved tools, anonymisation, low-sensitivity examples first.
- The durable outputs are the tested prompt, evaluation set, failure modes and user evidence.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run a chat prototype for one use case on 15 real examples. Record failure modes and write the evidence threshold required to move to the next rung.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.