AI Product Management: From Idea to Reliable AI FeaturesSpecs, prototypes and sourcing decisions · Lesson 7 of 16

Prototyping AI features with playgrounds and no-code tools

Article · 15 min · 8 min lecture

Video lecture

Prototyping AI features with playgrounds and no-code tools

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Prototyping AI features

  • Two months vs two days
  • Prototype behaviour first
  • A two-day loop to a decision

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Prototype the behaviour before you build the product

AI features have a unique risk: you do not know if the model can do the job well enough until you try it on real inputs. The fastest way to reduce that risk is to prototype the behaviour, not the interface, using prompt playgrounds, spreadsheets and no-code automation. A PM who can run a credible prototype in two days saves weeks of engineering on ideas that were never going to work.

Your prototyping toolkit (examples; features change)

Tool typeExamplesUse it to
Prompt playgrounds / consolesOpenAI Playground, Anthropic Console (Workbench), Google AI StudioIterate prompts, compare models, test structured output, save prompt versions
AI assistants with projects/knowledgeChatGPT projects or custom GPTs, Claude Projects, Gemini GemsQuick RAG-like prototypes over a few documents
Spreadsheet evaluationGoogle Sheets or Excel with AI functions or a small scriptRun a prompt over 50–200 rows and score results
No-code automationn8n, Zapier, MakeWire a prototype into email, forms, Slack or a CRM for a pilot
App/UI generatorsTools such as v0, Lovable, Bolt, or Claude artefactsClickable UI prototypes for user testing

Use company-approved tools and accounts for any real data; check each tool's data retention and training settings before uploading customer information.

The two-day prototype loop

Day 1: behaviour

  1. Collect 30–50 real inputs (anonymised if needed).
  2. Write a first prompt in a playground; include the output format you want.
  3. Run it over all inputs (spreadsheet or small script); do not cherry-pick.
  4. Score outputs quickly (pass/fail or 1–5), and note failure patterns.
  5. Iterate the prompt 3–5 times; try one stronger and one cheaper model.

Day 2: experience and value

  1. Put the best version behind a simple interface or automation (a form, a Slack bot, an n8n workflow).
  2. Put it in front of 3–5 real users doing real work (with supervision).
  3. Observe: where do they trust it, where do they check, where do they give up?
  4. Record time per task, edit rates and quotes.
  5. Decide: kill, iterate or write the PRD.

Wizard-of-Oz testing

When the AI is not ready but you want to test the experience, have a human secretly produce the "AI" output behind the interface. You learn whether users want the feature, what they ask, and which outputs they value, before investing in the model. Be ethical: debrief participants afterwards, and never deceive real customers in production.

Hands-on: a spreadsheet evaluation in 30 minutes

Set up a sheet with columns: input, expected, output_A, score_A, output_B, score_B, notes. Then run a prompt over every row with a short script against any OpenAI-compatible API (your company's approved provider):

# sheet_eval.py  (pip install openai pandas) - reads prompts from CSV, writes outputs for side-by-side scoring
import os
import pandas as pd
from openai import OpenAI

client = OpenAI(base_url=os.getenv("LLM_BASE_URL"), api_key=os.environ["LLM_API_KEY"])  # approved provider only
SYSTEM = open("prompt_v3.txt", encoding="utf-8").read()
df = pd.read_csv("inputs.csv")        # column: input

def run(model, text):
    try:
        r = client.chat.completions.create(model=model, temperature=0,
                                           messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": text}])
        return r.choices[0].message.content
    except Exception as e:
        return f"ERROR: {e}"

df["output_A"] = [run(os.environ["MODEL_A"], t) for t in df["input"]]
df["output_B"] = [run(os.environ["MODEL_B"], t) for t in df["input"]]
df.to_csv("results_for_scoring.csv", index=False)
print("Scored columns score_A / score_B are for reviewers to fill in.")

Score blind where possible (hide which model produced which column), then compute averages and read the worst ten.

Worked example: an Islamabad edtech tests "AI homework feedback"

The PM collected 40 real student essays (with consent), wrote a feedback prompt with a rubric, and ran it in a playground and a sheet. Teachers scored feedback quality: the first prompt was too generic; version 4, which quoted the student's own sentences and gave one priority fix, scored far better. A cheaper model was nearly as good for short essays but weaker for long ones. A Wizard-of-Oz pilot with five teachers showed they wanted feedback as editable comments inside their existing tool, not a separate app. The PRD that followed was grounded in evidence, including a cost comparison between the two models.

Pitfalls

  • Demoing the three best outputs instead of scoring all of them.
  • Prototyping only the happy path.
  • Uploading customer data into unapproved tools.
  • Treating a successful prototype as production-ready (no evals at scale, no security, no monitoring).

How to measure success

A prototype report: inputs used, prompt versions, model comparison, scores with failure patterns, user observations, time and edit measurements, and a clear kill/iterate/build recommendation.

Key takeaways

  • Prototype the behaviour on real inputs before building the product
  • Use playgrounds, projects, spreadsheets, no-code automation and UI generators
  • Run the two-day loop: behaviour on day one, experience and value on day two
  • Score every output, not just the best; compare a stronger and a cheaper model
  • Use approved tools for real data, and treat prototypes as evidence, not production

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What is the main risk that AI prototyping reduces?
  2. Why score all prototype outputs rather than showcasing the best?
  3. When is a Wizard-of-Oz test appropriate?

Put it into practice

Run the two-day prototype loop on one AI feature idea with 30–50 real inputs and 3–5 users, and write a one-page prototype report ending in kill, iterate or build.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.