---
title: "The AI platform landscape: providers, APIs and how to choose"
description: "Four ways to get model intelligence into your product 1. First-party model APIs : Anthropic (Claude API), OpenAI (API platform), Google (Gemini Developer…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/provider-landscape-and-choosing
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Platform foundations · lesson 1 of 19 · 14 min

# The AI platform landscape: providers, APIs and how to choose

## Four ways to get model intelligence into your product

1. **First-party model APIs**: Anthropic (Claude API), OpenAI (API platform), Google (Gemini Developer API). Newest models and features usually land here first.
2. **Cloud AI platforms**: Amazon Bedrock (plus Anthropic's Claude Platform on AWS), Microsoft Foundry (formerly Azure AI Foundry), and Google's Gemini Enterprise Agent Platform (formerly Vertex AI). You get enterprise contracts, IAM, private networking, regional hosting and consolidated billing, sometimes with a feature lag or subset.
3. **Open-weight models via hosted APIs**: Llama, Mistral, Qwen, DeepSeek, Gemma and OpenAI's gpt-oss models served by inference providers or cloud model catalogs, often through OpenAI-compatible endpoints.
4. **Self-hosting** open-weight models with vLLM, Ollama or similar on your own GPUs: maximum control, maximum operational burden.

Most real products combine two or three of these.

## What actually differs between providers

| Dimension | Why it matters |
|---|---|
| Model quality on **your** tasks | Public benchmarks rarely match your workload; run your own eval |
| Features | Tool use, structured outputs, vision, audio, long context, caching, batch, citations, server-side tools |
| Price | Per-token input/output, cached input, batch discounts; check live pricing pages |
| Latency and throughput | Time to first token, tokens per second, rate limits |
| Data terms | Training use, retention, zero-data-retention options, regional processing |
| Availability | Regions, cloud platforms, SLAs, status history |
| Ecosystem | SDK quality, agent tooling, MCP support, community |

## The current families (September 2026, verify before relying on this)

- **Anthropic Claude**: an Opus line for the hardest reasoning and agentic work, a Sonnet line balancing capability and cost, a Haiku line for fast, cheap tasks, plus a top-tier Fable line; 1M-token context on current top models; adaptive thinking with an effort control. Available on the Claude API, Claude Platform on AWS, Amazon Bedrock, Google's platform and Microsoft Foundry.
- **OpenAI**: GPT-5-series models via the Responses API (the recommended API) and the still-supported Chat Completions API, reasoning-effort controls, realtime and audio models, image and video generation, embeddings.
- **Google Gemini**: Pro and Flash tiers (with `-latest` aliases), thinking configuration, strong multimodal input (images, audio, video, PDFs), embeddings, image and video generation models; Gemini Developer API or the enterprise platform.

Model names and versions change every few months. **Treat model IDs as configuration, never as code constants**, and use each provider's models endpoint or docs to confirm what is available.

## A decision framework

1. **List your tasks** (e.g., classify leads, write product copy, extract invoice fields, answer support questions from docs).
2. **Define quality** per task with 30–100 real examples and a scoring method.
3. **Shortlist 2–3 models** across providers at different price points.
4. **Run the eval**; compare quality, latency, cost per completed task.
5. **Check constraints**: data residency, contracts, existing cloud, compliance.
6. **Design for change**: an abstraction layer (module 5) so switching is configuration, not a rewrite.

## Worked example: a Lahore e-commerce enabler

A company building storefronts for 300 small merchants needs: product descriptions in English and Urdu, customer-support replies grounded in each store's policies, and invoice extraction from photos.

- Eval results (illustrative): a mid-tier model from provider A wins on Urdu copy; provider B's small model is cheapest for extraction with acceptable accuracy; a larger model from provider A is needed for multi-turn support.
- Constraint: merchants' customer data must stay under agreements the company can show to banking partners, so they choose a cloud platform contract for support traffic.
- Architecture: one internal gateway with per-task routing; all three tasks share logging and cost tracking.

## Hands-on: a first call to each provider

```python
# pip install anthropic openai google-genai
import os
import anthropic
from openai import OpenAI
from google import genai

PROMPT = "Write a 20-word product description for a handmade Multani blue pottery mug."

claude = anthropic.Anthropic()                      # ANTHROPIC_API_KEY
msg = claude.messages.create(model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=300,
                             messages=[{"role": "user", "content": PROMPT}])
print("Claude:", "".join(b.text for b in msg.content if b.type == "text"))

oai = OpenAI()                                       # OPENAI_API_KEY
resp = oai.responses.create(model=os.environ.get("OPENAI_MODEL", "gpt-5.5"), input=PROMPT)
print("OpenAI:", resp.output_text)

gem = genai.Client()                                 # GEMINI_API_KEY (or GOOGLE_API_KEY)
g = gem.models.generate_content(model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"), contents=PROMPT)
print("Gemini:", g.text)
```

Default model IDs above come from each SDK's current documentation at the time of writing; override them via environment variables after checking the providers' model lists.

## Pitfalls

- Picking a provider from benchmark headlines instead of your own eval.
- Hard-coding model IDs, then scrambling when a model is retired.
- Ignoring data terms until legal review blocks the launch.
- Assuming every feature exists on every cloud platform version of a model.

## Measuring success

Per task: quality score, cost per completed task, p95 latency, and the time it takes you to switch a task to another model (target: a configuration change plus an eval run).

## Video lecture: The AI platform landscape: providers, APIs and how to choose

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. The AI platform landscape
2. Why it matters
3. Four access routes
4. Seven differences
5. Current families (verify!)
6. Engine vs car
7. Simple example: product blurbs
8. Business example: Lahore storefront builder
9. Six-step method
10. How big an eval?
11. Common mistakes
12. Keep a decision log
13. Deeper: the Lahore storefront choice (illustrative)
14. Watch me do it: first calls
15. Recap + try this now

## Lecture transcript

### The AI platform landscape

Picking an AI provider feels like picking a phone plan in a shop where the plans change every month. There are first party APIs, cloud platforms, open weight models and self hosting, and every vendor claims to be best. In this lesson, you'll get a clear map of the landscape and a decision method you can reuse every time a new model launches.

### Why it matters

Why does this matter? Because the provider decision touches everything downstream: quality, cost, speed, data privacy, and how painful it is to change your mind later. Think of it like choosing where to buy ingredients for a restaurant. The supplier's quality matters, but so do delivery times, contracts, food safety paperwork and whether you can switch suppliers if prices jump. The best restaurants don't marry a supplier; they design the kitchen so ingredients can change.

### Four access routes

There are four ways to get model intelligence into your product. First party APIs from Anthropic, OpenAI and Google, where new models and features usually land first. Cloud AI platforms: Amazon Bedrock, Microsoft Foundry, which used to be called Azure AI Foundry, and Google's Gemini Enterprise Agent Platform, which used to be Vertex AI. They add enterprise contracts, identity controls, private networking and regional hosting. Open weight models like Llama, Mistral, Qwen, DeepSeek, Gemma and OpenAI's gpt oss, served by hosted providers. And self hosting on your own GPUs.

### Seven differences

What actually differs between providers? Seven things. Quality on your tasks, which public benchmarks rarely predict. Features, like tool use, structured outputs, vision, audio, caching and batch. Price per token, cached input and batch discounts. Latency and rate limits. Data terms: training use, retention and zero retention options. Availability: regions, platforms and service levels. And the ecosystem around the SDKs. No provider wins all seven, which is why the answer is usually a mix.

### Current families (verify!)

A quick word on the current families, with a warning that names change every few months. Claude has Opus, Sonnet and Haiku lines plus a top tier Fable line, with very long context on current top models. OpenAI's GPT five series runs through the Responses API, its recommended interface, with the older Chat Completions API still supported. Gemini has Pro and Flash tiers with latest aliases and strong multimodal input. The durable lesson: treat model ids as configuration, never as constants in code.

### Engine vs car

Let's pause on the cloud platforms, because they confuse a lot of people. When you use Claude through Amazon Bedrock, or a GPT model through Microsoft Foundry, you're still using the same underlying model family, but the platform wraps it in its own identity system, billing, networking and regions. Here's the key idea: the model is the engine, and the platform is the car around it. The engine might be identical, but the dashboard, the insurance and where you're allowed to drive can differ. Sometimes a brand new feature reaches the first party API weeks before it reaches a cloud platform, so always check the platform's own feature list.

### Simple example: product blurbs

Here's a simple example. You want AI to write short product descriptions for a craft shop. You collect thirty real products, write what a good description looks like, and try three models: one from each provider. You score each output from one to five, note the cost per description and the time taken. Maybe the cheapest model scores four point two and the priciest scores four point four. For product blurbs, that small gain may not justify three times the price. The eval made the decision obvious.

### Business example: Lahore storefront builder

Now a realistic business example, with illustrative numbers. A company in Lahore builds online stores for three hundred small merchants. It needs product copy in English and Urdu, support replies grounded in each store's policies, and invoice data pulled from photos. Their eval shows one provider's mid tier model writes the best Urdu, another provider's small model is cheapest for invoice extraction, and a larger model handles multi turn support best. Because bank partners need contractual data controls, support traffic runs through a cloud platform contract. One internal gateway routes each task.

### Six-step method

The decision method in six steps. List your tasks. Define quality with thirty to a hundred real examples and a scoring method. Shortlist two or three models across providers and price points. Run the eval and compare quality, latency and cost per completed task. Check constraints like data residency, contracts and compliance. And design for change, with an abstraction layer so switching is a configuration change, not a rewrite. Repeat whenever a new model launches.

### How big an eval?

How big should your eval be? For a first decision, thirty to fifty examples per task is usually enough to spot big differences. If two models are within a few points, add more examples before you decide, or pick the cheaper one and keep watching. Include the awkward cases too: long inputs, mixed languages like English with Urdu or Arabic, messy formatting and questions the model should refuse or escalate. Those edge cases are where models actually differ, and they're exactly what a public benchmark never tests for your business.

### Common mistakes

Common mistakes. Choosing from benchmark headlines instead of your own data. Hard coding model ids, then scrambling when a model is retired. Leaving data terms until legal review blocks the launch. And assuming every feature exists on every cloud platform's version of a model. Features often arrive on first party APIs first, so check each platform's feature table before you commit.

### Keep a decision log

One more habit that separates professionals from hobbyists: keep a simple decision log. For each task, write the date, the models you compared, the eval scores, the cost per completed task, and why you chose what you chose. When a new model launches in three months, you rerun the same eval and add a line. When a stakeholder asks why you're using a particular provider, you have evidence instead of opinions. It takes five minutes per decision and saves hours of debate later.

### Deeper: the Lahore storefront choice (illustrative)

Let's deepen the Lahore storefront example with illustrative numbers. The company ran thirty products per merchant category through three models. For Urdu product copy, native speakers rated one provider's mid tier model clearly best. For invoice extraction from phone photos, a small model from another provider was accurate enough and several times cheaper per invoice. For multi turn support, only a larger model handled policy questions reliably. Instead of one provider for everything, they used three routes behind one gateway. Their monthly AI bill ended up lower than the single premium model plan they had first budgeted, and quality improved where customers actually noticed it: the Urdu copy.

### Watch me do it: first calls

Watch me do it. Let's walk through the first calls script. I import the three SDKs. The prompt asks for a twenty word description of a handmade Multani blue pottery mug. For Claude, I create the client, which reads the API key from the environment, and call messages create with the model from an environment variable, max tokens three hundred, and one user message. I join the text blocks and print them. For OpenAI, I create the client and call responses create with the model and the prompt as input, then print output text. For Gemini, I create the client, which reads the Gemini API key, and call models generate content with the model and contents, then print dot text. I run it: three short descriptions appear. Notice three different call shapes and three different ways to read the answer, but the same idea every time. That's exactly what we'll normalize later in the adapter lesson.

### Recap + try this now

Quick recap. There are four routes to model intelligence. Providers differ on seven dimensions, and none wins them all. Model ids are configuration. And your own eval, measured as cost per completed task, beats any leaderboard. Try this now: run the three first calls in the lesson's code, one per provider, using your own prompt. Then list three tasks from your product and pick two candidate models for each. That's the start of your eval plan.

## Key takeaways

- Model intelligence comes from first-party APIs, cloud platforms, hosted open-weight models or self-hosting.
- Providers differ on quality for your tasks, features, price, latency, data terms, availability and ecosystem.
- Treat model IDs as configuration and verify current models in each provider's docs.
- Choose with your own eval of cost per completed task, then check data and contract constraints.
- Design for change so switching providers is configuration, not a rewrite.

## Try it

List three AI tasks in your product, shortlist two models from different providers for each, and write the eval you would run to choose.

- [Next: API keys, authentication and secrets management](https://optimizeall.com/learn/ai-platform-apis-integration/api-keys-auth-and-secrets)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
