---
title: "Cloud AI platforms: Amazon Bedrock, Microsoft Foundry and…"
description: "Why enterprises use cloud AI platforms Large organizations often access models through their existing cloud rather than directly from model vendors: -…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/cloud-ai-platforms
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Cloud platforms and open-weight models · lesson 15 of 19 · 16 min

# Cloud AI platforms: Amazon Bedrock, Microsoft Foundry and Gemini Enterprise Agent Platform

## Why enterprises use cloud AI platforms

Large organizations often access models through their existing cloud rather than directly from model vendors:

- **Contracts and billing** through existing cloud agreements and marketplaces (and committed-spend programs).
- **Identity and access** with the cloud's IAM (roles, managed identities, service accounts) instead of static API keys.
- **Networking**: private endpoints, VPC/VNet integration, no public internet egress.
- **Data residency**: regional deployment options for regulated data.
- **Governance**: centralized logging, policies, guardrail services, and procurement already approved.
- **Model choice**: catalogs with many vendors' and open-weight models in one place.

Trade-offs: new models and features may arrive later than on first-party APIs, some features are unavailable, API shapes can differ, and prices and quotas are set by the platform.

## The three major platforms (September 2026; verify current names and availability)

**Amazon Bedrock (AWS)**
- Hosts models from Anthropic (Claude), Meta, Mistral, Amazon, OpenAI and others.
- APIs: the Bedrock Runtime (including the model-agnostic Converse API) and the newer "Mantle" endpoints that expose provider-native request shapes; the Anthropic SDK offers `AnthropicBedrockMantle` for Claude on Bedrock, and the OpenAI SDK offers a Bedrock provider for OpenAI models hosted there.
- Separate from Bedrock, Anthropic offers **Claude Platform on AWS**: Anthropic-operated access with AWS IAM and Marketplace billing and same-day feature parity with the Claude API.
- Related services: Bedrock Guardrails, Knowledge Bases, AgentCore for agents.

**Microsoft Foundry** (formerly Azure AI Foundry; renamed at Ignite in November 2025)
- Azure OpenAI models, plus Claude, Llama, Mistral and many others through its model catalog.
- Authentication with API keys or Microsoft Entra ID; the OpenAI SDK has an Azure client, and the Anthropic SDK has an `AnthropicFoundry` client.
- Integrates with Azure networking, Key Vault, monitoring and content safety services.

**Gemini Enterprise Agent Platform** (formerly Vertex AI; renamed in 2026, existing API endpoints continue to work)
- Gemini models plus partner and open models (including Claude) via its Model Garden.
- Authentication with Google Cloud IAM and Application Default Credentials; the Google Gen AI SDK switches with `genai.Client(enterprise=True, project=..., location=...)` (older code may use `vertexai=True`; check your SDK version); the Anthropic SDK offers `AnthropicVertex` for Claude.
- Batch prediction from Cloud Storage or BigQuery, explicit context caching, grounding and agent services.

## Hands-on: the same Claude call on three routes

```python
import os
from anthropic import Anthropic, AnthropicBedrockMantle, AnthropicVertex

prompt = [{"role": "user", "content": "Summarize our returns policy in two sentences: ..."}]

# 1) First-party Claude API (API key or workload identity)
first_party = Anthropic()
r1 = first_party.messages.create(model="claude-sonnet-5", max_tokens=300, messages=prompt)

# 2) Amazon Bedrock via the Mantle client (AWS credentials; model IDs carry an 'anthropic.' prefix)
bedrock = AnthropicBedrockMantle(aws_region=os.environ.get("AWS_REGION", "eu-west-2"))
r2 = bedrock.messages.create(model="anthropic.claude-sonnet-5", max_tokens=300, messages=prompt)

# 3) Google Cloud (pip install "anthropic[vertex]"; Application Default Credentials; region such as "global")
vertex = AnthropicVertex(project_id=os.environ["GCP_PROJECT"], region=os.environ.get("GCP_REGION", "global"))
r3 = vertex.messages.create(model="claude-sonnet-5", max_tokens=300, messages=prompt)
```

Model availability, IDs and supported features differ per platform and region; check each platform's model catalog and the provider's platform-availability notes before relying on a feature such as batch, server-side tools or the latest caching options.

## Choosing between direct and platform access

| Situation | Leaning |
|---|---|
| Startup, fastest access to new features | First-party APIs |
| Existing AWS/Azure/GCP enterprise agreement and IAM standards | That cloud's platform |
| Strict data residency or private networking | Platform with regional deployment and private endpoints |
| Multi-vendor experimentation with one bill | Platform model catalogs |
| Need a feature only on the first-party API | First-party API (possibly alongside the platform) |

Many companies use both: first-party APIs for innovation and prototypes, a cloud platform for regulated production workloads.

## Worked example: a UAE bank's customer-service assistant

Requirements: data processed in-region where available, private networking, bank-wide IAM, audit logs, and an approved vendor list. The bank deploys through its existing cloud platform contract, uses private endpoints and managed identities (no API keys), enables the platform's guardrail and logging services, and keeps a small first-party sandbox (with synthetic data only) for evaluating new models before they reach the platform. Compliance signs off faster because the platform is already on the approved list.

## Quotas and capacity planning

Platform quotas are set per model, per region and often per account, and defaults for newly enabled models can be modest. Before launch, estimate peak requests and tokens per minute (from your load tests or cost ledger), compare them with current quotas in the console, and request increases early. For predictable high volume, some platforms offer provisioned or reserved throughput options with different pricing; compare them against on-demand pricing using your real traffic profile.

## Pitfalls

- Assuming feature parity with first-party APIs.
- Hard-coding platform-specific model IDs throughout code.
- Forgetting quotas: platform quotas are often lower by default and require requests to raise.
- Mixing platforms without clear data-classification rules.

## Measuring success

Time from model release to availability in your approved platform, quota headroom, feature gaps affecting roadmap, and compliance review time for new AI use cases.

## Video lecture: Cloud AI platforms: Amazon Bedrock, Microsoft Foundry and Gemini Enterprise Agent Platform

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Cloud AI platforms
2. Why cloud platforms
3. The trade-offs
4. Amazon Bedrock
5. Microsoft Foundry
6. Gemini Enterprise Agent Platform
7. Check quotas early
8. Simple example: prototype → production
9. Choosing a route
10. Business example: UAE bank
11. Common mistakes
12. Keep a feature matrix
13. Deeper: the bank's second assistant
14. Watch me do it: one call, three routes
15. Recap + try this now

## Lecture transcript

### Cloud AI platforms

If you work in a bank, a hospital, a government agency or a large enterprise, there's a good chance you won't call AI models directly from their makers. You'll go through your company's cloud: Amazon, Microsoft or Google. In this lesson you'll learn why, what each major cloud AI platform offers, what changed with their recent renaming, and how to decide which route fits each workload.

### Why cloud platforms

Why do enterprises prefer cloud platforms? Contracts and billing through existing agreements. Identity and access through the cloud's own roles and managed identities, instead of static keys. Private networking, so traffic never crosses the public internet. Regional deployment for data residency. Centralized logging and governance. And a catalog of many vendors' models in one place. Think of buying through your company's approved supplier portal instead of a personal credit card. Slower sometimes, but procurement, security and finance are already happy.

### The trade-offs

The trade offs are real, though. New models and features may arrive later than on first party APIs, some features may be missing, request shapes can differ, and prices and quotas are set by the platform. So the question is never just which model. It's which model, on which platform, with which features, in which region.

### Amazon Bedrock

First, Amazon Bedrock. It hosts models from Anthropic, Meta, Mistral, Amazon, OpenAI and others. You can use its model agnostic Converse API, or newer Mantle endpoints that keep each provider's native request shape. The Anthropic SDK has a Bedrock Mantle client for Claude, and the OpenAI SDK has a Bedrock provider for OpenAI models hosted there. Separately, Anthropic offers Claude Platform on AWS, which is operated by Anthropic, uses AWS identity and billing, and matches the Claude API's features on the same day.

### Microsoft Foundry

Second, Microsoft Foundry. It used to be called Azure AI Foundry and was renamed at Microsoft Ignite in November twenty twenty five. It offers Azure OpenAI models, plus Claude, Llama, Mistral and many others through its catalog. You authenticate with keys or Microsoft Entra ID, and it integrates with Azure networking, Key Vault, monitoring and content safety. Both the OpenAI and Anthropic SDKs have dedicated clients for it.

### Gemini Enterprise Agent Platform

Third, Google's Gemini Enterprise Agent Platform, which was called Vertex AI until its renaming in twenty twenty six. Existing API endpoints keep working. It offers Gemini models plus partner and open models, including Claude, through its Model Garden. You authenticate with Google Cloud identity. The Google Gen AI SDK switches to it with an enterprise flag and your project and location, and the Anthropic SDK has a client for Claude on Google Cloud. It adds batch prediction from BigQuery or Cloud Storage, context caching and agent services.

### Check quotas early

Here's a practical tip that saves real frustration: check quotas early. On cloud platforms, default quotas for a newly enabled model are often modest, and raising them can take a request and some waiting. Teams build a feature, test it happily with a few users, then launch and hit throttling on day one. Before launch, estimate your peak requests and tokens per minute, compare with your current quota in the platform console, and file increase requests well ahead of time, per model and per region.

### Simple example: prototype → production

A simple example. A developer at a mid sized company already uses Claude through the first party API for prototypes. Their production team runs on Google Cloud. Moving the same feature into production is mostly a client swap: construct the Claude client for Google Cloud with the project and region, keep the same messages call, and confirm the model is available in that region. The lesson's code shows the same call on the first party API, on Bedrock and on Google Cloud side by side.

### Choosing a route

How do you choose? Startups wanting the newest features fastest usually lean to first party APIs. Organizations with existing cloud agreements and identity standards lean to that cloud's platform. Strict data residency or private networking points to a platform with regional deployment and private endpoints. Multi vendor experimentation with one bill suits platform catalogs. And if you need a feature only the first party API has, use it, alongside the platform where appropriate. Many companies run both.

### Business example: UAE bank

Now a realistic business example. A bank in the UAE wants a customer service assistant. Requirements: in region processing where available, private networking, bank wide identity, audit logs, and only approved vendors. It deploys through its existing cloud platform contract, with private endpoints and managed identities, so there are no API keys at all. It enables the platform's guardrail and logging services. And it keeps a small first party sandbox, using only synthetic data, to evaluate new models before they reach the platform. Compliance signs off faster because the platform is already approved.

### Common mistakes

Common mistakes. Assuming every first party feature exists on the platform. Hard coding platform specific model ids throughout your code. Forgetting quotas, which are often lower by default on platforms and need requests to raise. And mixing platforms without clear rules about which data class may go where.

### Keep a feature matrix

One more habit for platform users: keep a living feature matrix. For each model you use, list which platforms and regions it's on, and whether the features you rely on, like structured outputs, prompt caching, batch, citations or server side tools, are available there. Update it when providers announce changes. It sounds like paperwork, but it prevents the classic surprise where a feature works perfectly in the prototype on the first party API, then turns out to be missing in the production region.

### Deeper: the bank's second assistant

Let's deepen the UAE bank example. The bank's first AI assistant took months to approve, mainly because the security team had to assess a new vendor from scratch. For the second, they used their existing cloud platform, where the model provider was already on the approved list, identity came from managed identities, and traffic used private endpoints. The assessment focused only on the new use case: what data, which model, what guardrails. Approval took weeks rather than months, illustrative again. They still keep a first party sandbox with synthetic data for trying new models early, and move a model into production only once it appears in their platform region.

### Watch me do it: one call, three routes

Watch me do it. Let's walk through the same Claude call on three routes. The prompt is one user message asking for a two sentence policy summary. Route one, the first party API: the default client reads a key or workload identity, and I pass the bare model id. Route two, Amazon Bedrock through the Mantle client: I pass the AWS region, and the model id carries the anthropic dot prefix; credentials come from the AWS chain, so there's no Anthropic key. Route three, Google Cloud: the Vertex client takes my project id and a region such as global, uses application default credentials, and takes the bare model id. In every case, the messages create call is identical. That's the point: moving a workload between routes is mostly constructor and model id changes. Before switching, I check the platform's model list for my region and whether features I use, like batch or caching options, are available there.

### Recap + try this now

Quick recap. Cloud platforms wrap models in enterprise contracts, identity, networking, residency and governance. Bedrock, Microsoft Foundry and Gemini Enterprise Agent Platform all host many vendors' models, including Claude, and two of them were recently renamed. Features, ids and quotas differ, so verify, and consider using both first party and platform routes. Try this now: map your cloud agreements, residency needs and required features, decide which workloads go where, and make one test call through your chosen platform.

## Key takeaways

- Cloud AI platforms add contracts, IAM, private networking, residency and governance around models.
- Bedrock, Microsoft Foundry and Gemini Enterprise Agent Platform host many vendors' models, including Claude.
- Names changed recently: Azure AI Foundry is now Microsoft Foundry; Vertex AI is now Gemini Enterprise Agent Platform.
- Features, model IDs and quotas differ by platform and region; verify before relying on them.
- Many teams combine first-party APIs for innovation with a platform for regulated production.

## Try it

Map your organization's cloud agreements, data-residency needs and required features, then decide which workloads should use first-party APIs and which a cloud platform. Make one test call through your chosen platform.

- [Previous: Model routing: the right model for each request](https://optimizeall.com/learn/ai-platform-apis-integration/model-routing-strategies)
- [Next: Open-weight models via hosted APIs and self-hosting](https://optimizeall.com/learn/ai-platform-apis-integration/open-weight-models-via-apis)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
