---
title: "Prototyping AI features with playgrounds and no-code tools"
description: "Prototype the behaviour before you build the product AI features have a unique risk: you do not know if the model can do the job well enough until you…"
url: https://optimizeall.com/learn/ai-product-management/prototyping-with-playgrounds-and-no-code
updated: 2026-10-05
---

AI Product Management: From Idea to Reliable AI Features · Specs, prototypes and sourcing decisions · lesson 7 of 16 · 15 min

# Prototyping AI features with playgrounds and no-code tools

## Prototype the behaviour before you build the product

AI features have a unique risk: you do not know if the model can do the job well enough until you try it on real inputs. The fastest way to reduce that risk is to **prototype the behaviour**, not the interface, using prompt playgrounds, spreadsheets and no-code automation. A PM who can run a credible prototype in two days saves weeks of engineering on ideas that were never going to work.

## Your prototyping toolkit (examples; features change)

| Tool type | Examples | Use it to |
|---|---|---|
| **Prompt playgrounds / consoles** | OpenAI Playground, Anthropic Console (Workbench), Google AI Studio | Iterate prompts, compare models, test structured output, save prompt versions |
| **AI assistants with projects/knowledge** | ChatGPT projects or custom GPTs, Claude Projects, Gemini Gems | Quick RAG-like prototypes over a few documents |
| **Spreadsheet evaluation** | Google Sheets or Excel with AI functions or a small script | Run a prompt over 50–200 rows and score results |
| **No-code automation** | n8n, Zapier, Make | Wire a prototype into email, forms, Slack or a CRM for a pilot |
| **App/UI generators** | Tools such as v0, Lovable, Bolt, or Claude artefacts | Clickable UI prototypes for user testing |

Use company-approved tools and accounts for any real data; check each tool's data retention and training settings before uploading customer information.

## The two-day prototype loop

**Day 1: behaviour**

1. Collect 30–50 real inputs (anonymised if needed).
2. Write a first prompt in a playground; include the output format you want.
3. Run it over all inputs (spreadsheet or small script); do not cherry-pick.
4. Score outputs quickly (pass/fail or 1–5), and note failure patterns.
5. Iterate the prompt 3–5 times; try one stronger and one cheaper model.

**Day 2: experience and value**

6. Put the best version behind a simple interface or automation (a form, a Slack bot, an n8n workflow).
7. Put it in front of 3–5 real users doing real work (with supervision).
8. Observe: where do they trust it, where do they check, where do they give up?
9. Record time per task, edit rates and quotes.
10. Decide: kill, iterate or write the PRD.

## Wizard-of-Oz testing

When the AI is not ready but you want to test the experience, have a human secretly produce the "AI" output behind the interface. You learn whether users want the feature, what they ask, and which outputs they value, before investing in the model. Be ethical: debrief participants afterwards, and never deceive real customers in production.

## Hands-on: a spreadsheet evaluation in 30 minutes

Set up a sheet with columns: `input`, `expected`, `output_A`, `score_A`, `output_B`, `score_B`, `notes`. Then run a prompt over every row with a short script against any OpenAI-compatible API (your company's approved provider):

```python
# sheet_eval.py  (pip install openai pandas) - reads prompts from CSV, writes outputs for side-by-side scoring
import os
import pandas as pd
from openai import OpenAI

client = OpenAI(base_url=os.getenv("LLM_BASE_URL"), api_key=os.environ["LLM_API_KEY"])  # approved provider only
SYSTEM = open("prompt_v3.txt", encoding="utf-8").read()
df = pd.read_csv("inputs.csv")        # column: input

def run(model, text):
    try:
        r = client.chat.completions.create(model=model, temperature=0,
                                           messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": text}])
        return r.choices[0].message.content
    except Exception as e:
        return f"ERROR: {e}"

df["output_A"] = [run(os.environ["MODEL_A"], t) for t in df["input"]]
df["output_B"] = [run(os.environ["MODEL_B"], t) for t in df["input"]]
df.to_csv("results_for_scoring.csv", index=False)
print("Scored columns score_A / score_B are for reviewers to fill in.")
```

Score blind where possible (hide which model produced which column), then compute averages and read the worst ten.

## Worked example: an Islamabad edtech tests "AI homework feedback"

The PM collected 40 real student essays (with consent), wrote a feedback prompt with a rubric, and ran it in a playground and a sheet. Teachers scored feedback quality: the first prompt was too generic; version 4, which quoted the student's own sentences and gave one priority fix, scored far better. A cheaper model was nearly as good for short essays but weaker for long ones. A Wizard-of-Oz pilot with five teachers showed they wanted feedback as editable comments inside their existing tool, not a separate app. The PRD that followed was grounded in evidence, including a cost comparison between the two models.

## Pitfalls

- Demoing the three best outputs instead of scoring all of them.
- Prototyping only the happy path.
- Uploading customer data into unapproved tools.
- Treating a successful prototype as production-ready (no evals at scale, no security, no monitoring).

## How to measure success

A prototype report: inputs used, prompt versions, model comparison, scores with failure patterns, user observations, time and edit measurements, and a clear kill/iterate/build recommendation.

## Video lecture: Prototyping AI features with playgrounds and no-code tools

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Prototyping AI features
2. Why behaviour first
3. Toolkit
4. Day 1: behaviour
5. Day 2: experience and value
6. Wizard-of-Oz testing
7. Simple example: travel emails
8. Business example: Islamabad edtech
9. Hands-on: sheet_eval.py
10. Common mistakes
11. Another example: Dubai recruiters
12. When a prototype earns a PRD
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### Prototyping AI features

Here is a costly pattern. A team spends two months building a beautiful AI feature, launches it, and discovers the model simply cannot do the core task well enough on real data. The painful part is that they could have found out in two days. In this lesson you will learn to prototype the behaviour before you build the product, using playgrounds, spreadsheets and no-code tools, with a two-day loop that ends in a clear decision.

### Why behaviour first

Why prototype behaviour rather than screens? Because AI has a unique risk: you do not know if the model can do the job until you try it on real inputs. Here is an analogy. Before a restaurant builds a new kitchen, the chef cooks the new dish in the old kitchen and lets a few regulars taste it. You do not need the new kitchen to learn whether people like the dish. Playgrounds and spreadsheets are your old kitchen.

### Toolkit

Your toolkit. Prompt playgrounds and consoles, like OpenAI's Playground, Anthropic's Console and Google AI Studio, for iterating prompts and comparing models. AI assistants with projects or knowledge files for quick document-based prototypes. Spreadsheets, for running a prompt over fifty or two hundred rows and scoring results. No-code automation like n8n, Zapier or Make, to wire a prototype into email, forms or a CRM. And UI generators for clickable prototypes. One rule: use company-approved tools and check their data settings before any real customer data goes in.

### Day 1: behaviour

Day one is behaviour. Collect thirty to fifty real inputs, anonymised if needed. Write a first prompt with the exact output format you want. Run it over every input, not your favourite three. Score outputs quickly, pass or fail or one to five, and note the failure patterns. Iterate the prompt three to five times, and try one stronger and one cheaper model. By the end of the day you know whether the core behaviour is feasible, and roughly at what cost.

### Day 2: experience and value

Day two is experience and value. Put the best version behind something simple: a form, a chat bot in your team's messaging tool, or an n8n workflow. Put it in front of three to five real users doing real work, with supervision. Watch where they trust it, where they check it, and where they give up. Measure time per task and how much they edit. Then decide: kill it, iterate, or write the PRD.

### Wizard-of-Oz testing

One more technique: Wizard of Oz. When the AI is not ready but you want to test the experience, a person secretly produces the AI output behind the interface. You learn whether users want the feature and what they ask for, before investing in the model. Be ethical about it: debrief participants afterwards, and never deceive real customers in production.

### Simple example: travel emails

A simple example. A PM at a small travel agency wonders whether AI can turn messy customer emails into structured booking requests. She pastes thirty real, anonymised emails into a spreadsheet, writes one prompt that outputs destination, dates, travelers and budget, runs it, and scores each row. Twenty-five are perfect, three miss the budget, two misread dates written in a local format. One prompt tweak about date formats fixes both. Feasibility answered in an afternoon.

### Business example: Islamabad edtech

Now a realistic business example. An Islamabad edtech company explored AI feedback on student essays. The PM collected forty real essays with consent, wrote a rubric-based prompt, and ran it in a playground and a sheet. Teachers scored the feedback. Version one was generic. Version four, which quoted the student's own sentences and gave one priority fix, scored far better. A cheaper model was nearly as good on short essays but weaker on long ones. A Wizard-of-Oz pilot with five teachers revealed they wanted feedback as editable comments inside their existing tool, not a new app. The PRD that followed was grounded in evidence.

### Hands-on: sheet_eval.py

The lesson includes a small script for spreadsheet evaluation. It reads your inputs from a CSV file, runs the same system prompt through two models on any OpenAI-compatible provider your company approves, and writes both outputs side by side for reviewers to score. Hide which model produced which column when scoring, then compute averages and read the worst ten outputs. The worst ten teach you more than the average.

### Common mistakes

Common mistakes. Demoing the three best outputs and hiding the rest. Prototyping only the happy path. Uploading customer data into tools your company has not approved. And treating a successful prototype as production-ready, when it has no evaluation at scale, no security review and no monitoring. A prototype is evidence, not a product.

### Another example: Dubai recruiters

Another quick example. A recruitment agency in Dubai wonders whether AI can write first-draft job descriptions from a hiring manager's rough notes. The PM builds an n8n workflow in an afternoon: a form collects the notes, a model drafts the description in the agency's template, and the draft lands in a shared folder. Five recruiters use it for a week. They love the structure, but the AI keeps inventing salary ranges. One prompt rule and a required salary field in the form fix it.

### When a prototype earns a PRD

How do you know when a prototype has earned a PRD? Three signals. The behaviour passes on most real inputs, with failure patterns you understand and believe you can fix. Real users came back to it without being asked. And the numbers, time saved and cost per task, make sense at your expected volume. If you have all three, write the PRD. If you have none, kill it without regret. If you have one or two, iterate once more.

### Try this now

Try this now. Pick one AI feature idea. Gather thirty real inputs this week, anonymised, and run one prompt over all of them in a spreadsheet. Score every output. Then write a half-page note: the pass rate, the top two failure patterns, the cost per task you observed, and your recommendation to kill, iterate or build.

### Watch me do it

Watch me do it. Idea: AI turns messy supplier quotes into a comparison table. Day one, morning: I collect thirty real quotes, anonymised, and paste them into a spreadsheet. I write a prompt in a playground asking for supplier, item, unit price, currency, delivery days and validity. I run the sheet script over all thirty with two models and score each field. Stronger model: twenty-six fully correct. Cheaper model: twenty-one. Failures cluster around quotes in dirhams and riyals without a currency symbol. I add one instruction about inferring currency from the supplier's country, and the cheaper model rises to twenty-five. Afternoon: I wire the cheaper model into an n8n workflow, from an email inbox to a shared sheet. Day two: three buyers use it on real quotes while I watch. They trust it but want the original quote line shown beside each price. Recommendation: build, with source-line display as a requirement.

### Recap

Recap. Prototype the behaviour on real inputs before building the product. Use playgrounds, spreadsheets, no-code tools and UI generators, with approved accounts. Run the two-day loop, score every output, compare a stronger and a cheaper model, and watch real users. End with a clear decision. In the next lesson, we will decide whether to build, buy or call an API.

## Key takeaways

- Prototype the behaviour on real inputs before building the product
- Use playgrounds, projects, spreadsheets, no-code automation and UI generators
- Run the two-day loop: behaviour on day one, experience and value on day two
- Score every output, not just the best; compare a stronger and a cheaper model
- Use approved tools for real data, and treat prototypes as evidence, not production

## Try it

Run the two-day prototype loop on one AI feature idea with 30–50 real inputs and 3–5 users, and write a one-page prototype report ending in kill, iterate or build.

- [Previous: Writing AI PRDs with evaluation criteria](https://optimizeall.com/learn/ai-product-management/writing-ai-prds)
- [Next: Build, buy or API: data and model choices for PMs](https://optimizeall.com/learn/ai-product-management/build-buy-or-api)
- [All lessons of AI Product Management: From Idea to Reliable AI Features](https://optimizeall.com/learn/ai-product-management)
