Skip to content

AI Product Management: From Idea to Reliable AI Features · Specs, prototypes and sourcing decisions · lesson 8 of 16 · 15 min

Build, buy or API: data and model choices for PMs

Four ways to get AI capability

| Option | What it means | Speed | Control | Differentiation | |---|---|---|---|---| | Buy | Use AI features inside existing SaaS (CRM, helpdesk, office suite) | Fastest | Low | Low | | API | Call foundation models (closed or open-weight hosted) from your product | Fast | Medium | Medium to high (via workflow, data, UX) | | Build on open weights | Self-host open-weight models, possibly fine-tuned | Slower | High | High where data/privacy matter | | Train from scratch | Pre-train your own foundation model | Very slow, very expensive | Full | Rarely justified outside a few organisations |

For most product teams the real decision is buy vs API vs open weights, often mixed per feature.

Decision criteria

  1. Is this capability core to your differentiation? If not, buy it. (An internal meeting summariser is rarely your moat.)
  2. Data sensitivity and residency: can the data leave your boundary? Which regions are allowed? Enterprise API terms (no training on your data, retention limits, regional processing) may satisfy requirements; if not, consider self-hosted open weights.
  3. Quality needs: does the task need frontier capability, or does a smaller model pass your golden set?
  4. Cost at scale: API per-token cost vs infrastructure and people for self-hosting; model the volume.
  5. Latency and availability: real-time voice or on-device needs may push toward smaller or local models.
  6. Team capability: do you have engineers to operate models, evals and monitoring?
  7. Vendor risk: pricing changes, model retirements, offering changes (for example, hosted fine-tuning being wound down), outages. Mitigate with abstraction and tested alternatives.

Model selection is a portfolio, not a marriage

Mature AI products use multiple models: a strong model for hard reasoning, a fast cheap model for classification and routing, an embedding model for search, maybe a speech or vision model. Routing requests to the right model can cut cost substantially while improving quality where it matters. Your architecture should make swapping a model a configuration change, validated by your golden set.

Data choices: the PM's questions

  • What data does the AI need? (documents, CRM fields, product catalogue, conversation history)
  • Who owns it, and may we use it for this purpose? (privacy notices, contracts, consent)
  • How fresh must it be? (real-time feed vs nightly sync)
  • How good is it? (duplicates, outdated policies, missing Arabic versions)
  • What must be excluded? (sensitive fields, other customers' data)

Often the biggest AI quality win is cleaning the knowledge base, not changing the model.

Hands-on: a build/buy/API decision record

## Decision record: [capability], [date]
**Capability:** Arabic + English customer support reply drafting
**Options considered:**
| Option | Quality on golden set | Cost/1k tasks (illustrative) | Data/residency | Time to launch | Vendor risk |
|---|---|---|---|---|---|
| Helpdesk vendor's built-in AI | 3.6/5 | included in plan | vendor region, no custom data | 1 week | medium |
| API: strong model + our RAG | 4.4/5 | $X | enterprise terms, regional processing | 6 weeks | medium (mitigated by abstraction) |
| Self-hosted open-weight + RAG | 4.1/5 | $Y + ops | in-country | 10 weeks | low vendor / higher ops |
**Decision:** API + our RAG now; re-evaluate self-hosting at 3x volume or if residency rules tighten.
**Reversibility:** model abstraction layer; golden set v1 for switching tests.
**Owner / review date:** [name], [date + 6 months]

Worked example: three features, three answers at a Dubai fintech

  • Internal meeting notes: buy (the office suite's built-in AI meets the need; not differentiating).
  • Customer support drafting: API with the company's own retrieval over policies; enterprise terms with regional processing met compliance; strong quality needed for Arabic.
  • Transaction categorisation for millions of events: a small open-weight model fine-tuned and self-hosted, because volume made per-token API costs high and latency mattered; evaluated against the API baseline first.

Pitfalls

  • Building what you could buy, for non-differentiating capabilities.
  • Buying a black box for your core differentiator.
  • Choosing a model before defining the golden set.
  • Ignoring data quality and permissions while debating models.

How to measure success

Each AI capability has a decision record with options, evidence from your golden set, cost at expected volume, data and residency analysis, reversibility plan and a review date.

Video lecture: Build, buy or API: data and model choices for PMs

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

  1. Build, buy or API
  2. Analogy: a restaurant's bread
  3. Four options
  4. Seven criteria
  5. Models as a portfolio
  6. Data questions
  7. Simple example: accounting firm
  8. Business example: Dubai fintech
  9. Hands-on: decision record
  10. Common mistakes
  11. Another example: UK retailer
  12. Vendor terms checklist
  13. Try this now
  14. Watch me do it
  15. Recap

Lecture transcript

Build, buy or API

Every AI product team faces the same fork in the road, usually more than once a quarter. Do we buy an AI feature from a vendor, call a model through an API, or build on open-weight models ourselves? Get it wrong one way and you waste months building a commodity. Get it wrong the other way and you hand your core differentiator to a black box. In this lesson you will learn a clear decision method, how to think about models as a portfolio, and the data questions PMs must ask.

Analogy: a restaurant's bread

Here is an analogy. Think about how a restaurant gets its bread. It can buy bread from a bakery: fast, reliable, but the same bread other restaurants serve. It can buy dough and bake it in-house: more control over the final product. Or it can mill its own flour: full control, very expensive, rarely worth it. Buying AI features is buying bread. Using an API is baking with bought dough. Building on open weights is milling your own flour for the one bread you are famous for.

Four options

So the options are four. Buy: AI inside existing software, fastest, least control, least differentiation. API: call foundation models from your product, fast, with differentiation coming from your workflow, data and UX. Build on open weights: self-host models, possibly fine-tuned, slower but with high control, especially where privacy matters. And train from scratch, which is rarely justified outside a handful of organisations. For most teams, the real choice is buy, API or open weights, and it is often different per feature.

Seven criteria

Seven criteria decide it. Is the capability core to your differentiation? If not, buy. Can the data leave your boundary, and in which regions? Does the task need frontier quality, or does a smaller model pass your golden set? What is the cost at your real volume? Do latency or on-device needs push you to smaller or local models? Does your team have the skills to operate models? And what is the vendor risk: price changes, model retirements, even whole offerings being wound down?

Models as a portfolio

Next, stop thinking of model choice as a marriage. Mature AI products run a portfolio: a strong model for hard reasoning, a fast cheap model for classification and routing, an embedding model for search, maybe speech or vision models. Routing each request to the right model can cut cost substantially while improving quality where it matters. Your architecture should make swapping a model a configuration change, and your golden set should validate every swap.

Data questions

Now the data questions every PM must ask, because they often matter more than the model. What data does the AI need? Who owns it, and may we use it for this purpose? How fresh must it be? How good is it: duplicates, outdated policies, missing Arabic versions? And what must be excluded? In many projects, the biggest quality jump comes from cleaning the knowledge base, not from switching models.

Simple example: accounting firm

A simple example. A small accounting firm wants three AI capabilities: meeting notes, client email drafting and invoice data extraction. Meeting notes are not differentiating, and their office suite already has it: buy. Client emails need their tone and templates but no special infrastructure: an API model with their templates. Invoice extraction runs on thousands of scanned invoices a month: they compare an API model and a small specialised model on fifty invoices, and pick whichever passes at lower cost. Three features, three different answers.

Business example: Dubai fintech

Now a realistic business example at a Dubai fintech. Internal meeting notes: they buy, using their office suite's built-in AI. Customer support drafting: an API model with the company's own retrieval over policies, under enterprise terms with regional processing that met compliance, because Arabic quality needed a strong model. Transaction categorisation, millions of events a month with tight latency: a small open-weight model fine-tuned and self-hosted, evaluated against the API baseline first, because per-token costs at that volume were too high.

Hands-on: decision record

Capture every choice in a decision record. The capability. The options you considered, each scored on your golden set, with cost per thousand tasks, data and residency implications, time to launch and vendor risk. The decision. How you would reverse it: an abstraction layer and the golden set for switching tests. And an owner with a review date, typically six months out, or earlier if volume or regulation changes.

Common mistakes

Common mistakes. Building what you could buy for capabilities that do not differentiate you. Buying a black box for your core differentiator. Choosing a model before you have a golden set to judge it. And debating models for weeks while the knowledge base is full of duplicates and outdated policies.

Another example: UK retailer

Another example, from a UK retailer. They wanted product recommendations. The e-commerce platform they use offers built-in AI recommendations, which performed reasonably. But their differentiator is expert styling advice. So they bought the built-in recommendations for the long tail of products, and built a styling assistant on an API model with their stylists' guides and outfit rules, where their brand actually stands apart. Buy the commodity, build the signature.

Vendor terms checklist

A note on vendor terms, because PMs often overlook them. Before choosing an API or a bought feature, check whether the vendor may use your data to train its models, how long it keeps prompts and outputs, which regions process the data, which sub-processors are involved, and what happens if the vendor retires the model you depend on. Put those answers in your decision record. They matter as much as the benchmark scores.

Try this now

Try this now. List every AI capability in your product roadmap, even the small ones. Mark each as differentiating or not. For every one that is not differentiating, write the name of a tool you could buy. For the differentiating ones, write the single biggest data question you have not answered yet. That list will reshape your next planning meeting.

Watch me do it

Watch me do it. I'm writing a decision record for AI invoice extraction at a mid-size distributor. Option one, buy: their accounting software's built-in extraction. I test it on fifty real invoices: thirty-eight fully correct, weak on Arabic invoices. Option two, API: a strong model with our prompt and schema, forty-six correct, at a cost per invoice I compute from token counts. Option three, open weights: a smaller self-hosted vision-capable model, forty-one correct, cheaper at volume but needing an engineer to run it. Then data questions: invoices contain supplier bank details, so I check the API vendor's terms, regional processing and retention; they meet our policy. Decision: API now, because quality matters and volume is moderate. Reversibility: an abstraction layer and this fifty-invoice golden set. Review date: in six months, or if volume triples. The record is one page and legal signs it the same day.

Recap

Recap. Choose between buying, calling an API and building on open weights, feature by feature. Decide with seven criteria, treat models as a portfolio, make swaps a configuration change validated by your golden set, and ask hard data questions early. Record each decision with a way to reverse it. Next module: the economics, from tokens to cost per task and pricing.

Key takeaways

  • Options: buy, API, build on open weights, or (rarely) train from scratch
  • Decide by differentiation, data sensitivity, quality, cost at scale, latency, team capability and vendor risk
  • Treat models as a portfolio; make swapping a configuration change validated by the golden set
  • Data rights, freshness and quality often matter more than the model
  • Write a decision record with a reversibility plan and review date

Try it

Write a build/buy/API decision record for one AI capability, with at least three options scored on your golden set, cost, data/residency and reversibility.