Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APIPlatform foundations · Lesson 1 of 19

The AI platform landscape: providers, APIs and how to choose

Article · 14 min · 9 min lecture

Video lecture

The AI platform landscape: providers, APIs and how to choose

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

The AI platform landscape

  • Four ways to access models
  • What really differs
  • A reusable decision method

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Four ways to get model intelligence into your product

  1. First-party model APIs: Anthropic (Claude API), OpenAI (API platform), Google (Gemini Developer API). Newest models and features usually land here first.
  2. Cloud AI platforms: Amazon Bedrock (plus Anthropic's Claude Platform on AWS), Microsoft Foundry (formerly Azure AI Foundry), and Google's Gemini Enterprise Agent Platform (formerly Vertex AI). You get enterprise contracts, IAM, private networking, regional hosting and consolidated billing, sometimes with a feature lag or subset.
  3. Open-weight models via hosted APIs: Llama, Mistral, Qwen, DeepSeek, Gemma and OpenAI's gpt-oss models served by inference providers or cloud model catalogs, often through OpenAI-compatible endpoints.
  4. Self-hosting open-weight models with vLLM, Ollama or similar on your own GPUs: maximum control, maximum operational burden.

Most real products combine two or three of these.

What actually differs between providers

DimensionWhy it matters
Model quality on your tasksPublic benchmarks rarely match your workload; run your own eval
FeaturesTool use, structured outputs, vision, audio, long context, caching, batch, citations, server-side tools
PricePer-token input/output, cached input, batch discounts; check live pricing pages
Latency and throughputTime to first token, tokens per second, rate limits
Data termsTraining use, retention, zero-data-retention options, regional processing
AvailabilityRegions, cloud platforms, SLAs, status history
EcosystemSDK quality, agent tooling, MCP support, community

The current families (September 2026, verify before relying on this)

  • Anthropic Claude: an Opus line for the hardest reasoning and agentic work, a Sonnet line balancing capability and cost, a Haiku line for fast, cheap tasks, plus a top-tier Fable line; 1M-token context on current top models; adaptive thinking with an effort control. Available on the Claude API, Claude Platform on AWS, Amazon Bedrock, Google's platform and Microsoft Foundry.
  • OpenAI: GPT-5-series models via the Responses API (the recommended API) and the still-supported Chat Completions API, reasoning-effort controls, realtime and audio models, image and video generation, embeddings.
  • Google Gemini: Pro and Flash tiers (with -latest aliases), thinking configuration, strong multimodal input (images, audio, video, PDFs), embeddings, image and video generation models; Gemini Developer API or the enterprise platform.

Model names and versions change every few months. Treat model IDs as configuration, never as code constants, and use each provider's models endpoint or docs to confirm what is available.

A decision framework

  1. List your tasks (e.g., classify leads, write product copy, extract invoice fields, answer support questions from docs).
  2. Define quality per task with 30–100 real examples and a scoring method.
  3. Shortlist 2–3 models across providers at different price points.
  4. Run the eval; compare quality, latency, cost per completed task.
  5. Check constraints: data residency, contracts, existing cloud, compliance.
  6. Design for change: an abstraction layer (module 5) so switching is configuration, not a rewrite.

Worked example: a Lahore e-commerce enabler

A company building storefronts for 300 small merchants needs: product descriptions in English and Urdu, customer-support replies grounded in each store's policies, and invoice extraction from photos.

  • Eval results (illustrative): a mid-tier model from provider A wins on Urdu copy; provider B's small model is cheapest for extraction with acceptable accuracy; a larger model from provider A is needed for multi-turn support.
  • Constraint: merchants' customer data must stay under agreements the company can show to banking partners, so they choose a cloud platform contract for support traffic.
  • Architecture: one internal gateway with per-task routing; all three tasks share logging and cost tracking.

Hands-on: a first call to each provider

# pip install anthropic openai google-genai
import os
import anthropic
from openai import OpenAI
from google import genai

PROMPT = "Write a 20-word product description for a handmade Multani blue pottery mug."

claude = anthropic.Anthropic()                      # ANTHROPIC_API_KEY
msg = claude.messages.create(model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=300,
                             messages=[{"role": "user", "content": PROMPT}])
print("Claude:", "".join(b.text for b in msg.content if b.type == "text"))

oai = OpenAI()                                       # OPENAI_API_KEY
resp = oai.responses.create(model=os.environ.get("OPENAI_MODEL", "gpt-5.5"), input=PROMPT)
print("OpenAI:", resp.output_text)

gem = genai.Client()                                 # GEMINI_API_KEY (or GOOGLE_API_KEY)
g = gem.models.generate_content(model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"), contents=PROMPT)
print("Gemini:", g.text)

Default model IDs above come from each SDK's current documentation at the time of writing; override them via environment variables after checking the providers' model lists.

Pitfalls

  • Picking a provider from benchmark headlines instead of your own eval.
  • Hard-coding model IDs, then scrambling when a model is retired.
  • Ignoring data terms until legal review blocks the launch.
  • Assuming every feature exists on every cloud platform version of a model.

Measuring success

Per task: quality score, cost per completed task, p95 latency, and the time it takes you to switch a task to another model (target: a configuration change plus an eval run).

Key takeaways

  • Model intelligence comes from first-party APIs, cloud platforms, hosted open-weight models or self-hosting.
  • Providers differ on quality for your tasks, features, price, latency, data terms, availability and ecosystem.
  • Treat model IDs as configuration and verify current models in each provider's docs.
  • Choose with your own eval of cost per completed task, then check data and contract constraints.
  • Design for change so switching providers is configuration, not a rewrite.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What is the most reliable way to choose a model for invoice extraction?
  2. Why treat model IDs as configuration?
  3. A bank partner requires contractual data controls. Which option often helps most?

Put it into practice

List three AI tasks in your product, shortlist two models from different providers for each, and write the eval you would run to choose.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.