Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APICloud platforms and open-weight models · Lesson 15 of 19

Cloud AI platforms: Amazon Bedrock, Microsoft Foundry and Gemini Enterprise Agent Platform

Article · 16 min · 9 min lecture

Video lecture

Cloud AI platforms: Amazon Bedrock, Microsoft Foundry and Gemini Enterprise Agent Platform

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Cloud AI platforms

  • Why enterprises use them
  • Bedrock, Microsoft Foundry, Gemini Enterprise Agent Platform
  • Choosing the route

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why enterprises use cloud AI platforms

Large organizations often access models through their existing cloud rather than directly from model vendors:

  • Contracts and billing through existing cloud agreements and marketplaces (and committed-spend programs).
  • Identity and access with the cloud's IAM (roles, managed identities, service accounts) instead of static API keys.
  • Networking: private endpoints, VPC/VNet integration, no public internet egress.
  • Data residency: regional deployment options for regulated data.
  • Governance: centralized logging, policies, guardrail services, and procurement already approved.
  • Model choice: catalogs with many vendors' and open-weight models in one place.

Trade-offs: new models and features may arrive later than on first-party APIs, some features are unavailable, API shapes can differ, and prices and quotas are set by the platform.

The three major platforms (September 2026; verify current names and availability)

Amazon Bedrock (AWS)

  • Hosts models from Anthropic (Claude), Meta, Mistral, Amazon, OpenAI and others.
  • APIs: the Bedrock Runtime (including the model-agnostic Converse API) and the newer "Mantle" endpoints that expose provider-native request shapes; the Anthropic SDK offers AnthropicBedrockMantle for Claude on Bedrock, and the OpenAI SDK offers a Bedrock provider for OpenAI models hosted there.
  • Separate from Bedrock, Anthropic offers Claude Platform on AWS: Anthropic-operated access with AWS IAM and Marketplace billing and same-day feature parity with the Claude API.
  • Related services: Bedrock Guardrails, Knowledge Bases, AgentCore for agents.

Microsoft Foundry (formerly Azure AI Foundry; renamed at Ignite in November 2025)

  • Azure OpenAI models, plus Claude, Llama, Mistral and many others through its model catalog.
  • Authentication with API keys or Microsoft Entra ID; the OpenAI SDK has an Azure client, and the Anthropic SDK has an AnthropicFoundry client.
  • Integrates with Azure networking, Key Vault, monitoring and content safety services.

Gemini Enterprise Agent Platform (formerly Vertex AI; renamed in 2026, existing API endpoints continue to work)

  • Gemini models plus partner and open models (including Claude) via its Model Garden.
  • Authentication with Google Cloud IAM and Application Default Credentials; the Google Gen AI SDK switches with genai.Client(enterprise=True, project=..., location=...) (older code may use vertexai=True; check your SDK version); the Anthropic SDK offers AnthropicVertex for Claude.
  • Batch prediction from Cloud Storage or BigQuery, explicit context caching, grounding and agent services.

Hands-on: the same Claude call on three routes

import os
from anthropic import Anthropic, AnthropicBedrockMantle, AnthropicVertex

prompt = [{"role": "user", "content": "Summarize our returns policy in two sentences: ..."}]

# 1) First-party Claude API (API key or workload identity)
first_party = Anthropic()
r1 = first_party.messages.create(model="claude-sonnet-5", max_tokens=300, messages=prompt)

# 2) Amazon Bedrock via the Mantle client (AWS credentials; model IDs carry an 'anthropic.' prefix)
bedrock = AnthropicBedrockMantle(aws_region=os.environ.get("AWS_REGION", "eu-west-2"))
r2 = bedrock.messages.create(model="anthropic.claude-sonnet-5", max_tokens=300, messages=prompt)

# 3) Google Cloud (pip install "anthropic[vertex]"; Application Default Credentials; region such as "global")
vertex = AnthropicVertex(project_id=os.environ["GCP_PROJECT"], region=os.environ.get("GCP_REGION", "global"))
r3 = vertex.messages.create(model="claude-sonnet-5", max_tokens=300, messages=prompt)

Model availability, IDs and supported features differ per platform and region; check each platform's model catalog and the provider's platform-availability notes before relying on a feature such as batch, server-side tools or the latest caching options.

Choosing between direct and platform access

SituationLeaning
Startup, fastest access to new featuresFirst-party APIs
Existing AWS/Azure/GCP enterprise agreement and IAM standardsThat cloud's platform
Strict data residency or private networkingPlatform with regional deployment and private endpoints
Multi-vendor experimentation with one billPlatform model catalogs
Need a feature only on the first-party APIFirst-party API (possibly alongside the platform)

Many companies use both: first-party APIs for innovation and prototypes, a cloud platform for regulated production workloads.

Worked example: a UAE bank's customer-service assistant

Requirements: data processed in-region where available, private networking, bank-wide IAM, audit logs, and an approved vendor list. The bank deploys through its existing cloud platform contract, uses private endpoints and managed identities (no API keys), enables the platform's guardrail and logging services, and keeps a small first-party sandbox (with synthetic data only) for evaluating new models before they reach the platform. Compliance signs off faster because the platform is already on the approved list.

Quotas and capacity planning

Platform quotas are set per model, per region and often per account, and defaults for newly enabled models can be modest. Before launch, estimate peak requests and tokens per minute (from your load tests or cost ledger), compare them with current quotas in the console, and request increases early. For predictable high volume, some platforms offer provisioned or reserved throughput options with different pricing; compare them against on-demand pricing using your real traffic profile.

Pitfalls

  • Assuming feature parity with first-party APIs.
  • Hard-coding platform-specific model IDs throughout code.
  • Forgetting quotas: platform quotas are often lower by default and require requests to raise.
  • Mixing platforms without clear data-classification rules.

Measuring success

Time from model release to availability in your approved platform, quota headroom, feature gaps affecting roadmap, and compliance review time for new AI use cases.

Key takeaways

  • Cloud AI platforms add contracts, IAM, private networking, residency and governance around models.
  • Bedrock, Microsoft Foundry and Gemini Enterprise Agent Platform host many vendors' models, including Claude.
  • Names changed recently: Azure AI Foundry is now Microsoft Foundry; Vertex AI is now Gemini Enterprise Agent Platform.
  • Features, model IDs and quotas differ by platform and region; verify before relying on them.
  • Many teams combine first-party APIs for innovation with a platform for regulated production.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A bank requires managed identities, private endpoints and in-region processing. Which option usually fits best?
  2. What is Microsoft Foundry?
  3. Your roadmap depends on a feature only on the first-party API. What is a reasonable approach?

Put it into practice

Map your organization's cloud agreements, data-residency needs and required features, then decide which workloads should use first-party APIs and which a cloud platform. Make one test call through your chosen platform.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.