---
title: "Hosted fine-tuning: which platforms, which models, which…"
description: "A fast-moving market Hosted fine-tuning lets you upload data, run a job and call a custom model without managing GPUs. It is also one of the…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/hosted-fine-tuning-apis
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Evaluation, hosted platforms, embeddings and cost · lesson 12 of 16 · 15 min

# Hosted fine-tuning: which platforms, which models, which trade-offs

## A fast-moving market

Hosted fine-tuning lets you upload data, run a job and call a custom model without managing GPUs. It is also one of the fastest-changing parts of the AI market. The snapshot below reflects official documentation as of September 2026; **verify current models, methods, regions and pricing before you commit**, and design so you can move.

## The landscape (September 2026 snapshot)

| Provider | What is offered (verify) | Notes |
|---|---|---|
| **Google Cloud Vertex AI** (Google's docs now also brand it the Gemini Enterprise Agent Platform) | Supervised fine-tuning for Gemini 2.5 Pro, Gemini 2.5 Flash and Gemini 2.5 Flash-Lite (text, image, audio, video and document data per docs); preference tuning for Gemini; supervised and distillation tuning for selected open models | Via the Google Gen AI SDK or REST; regional availability matters for residency |
| **Amazon Bedrock** | Model customization including supervised fine-tuning, distillation and reinforcement fine-tuning for Amazon Nova models (RFT starting with Nova 2 Lite); other base models vary by region | Graders for RFT via AWS Lambda or model-as-judge; check whether custom models need dedicated capacity for inference |
| **Microsoft Foundry (Azure)** | Fine-tuning for GPT-4.1, GPT-4.1-mini, GPT-4.1-nano, GPT-4o and GPT-4o-mini with supervised methods, plus other methods and models with varying (sometimes gated) access | Microsoft documentation noted GPT-5-family models were not supported for fine-tuning at the time of writing |
| **OpenAI** | Self-serve fine-tuning (SFT, DPO, RFT, vision) historically | In 2026 OpenAI began winding down its self-serve fine-tuning platform: closed to new organizations, with job creation for existing customers being phased out; existing fine-tuned models remain available for inference until their base models are deprecated |
| **Mistral AI** | Has offered a fine-tuning API for its models | Its docs now place the self-serve fine-tuning API under legacy resources; check current status |
| **Inference platforms for open models** | Several providers offer LoRA/SFT (and sometimes DPO) on open-weight models with serving included | You can often download the resulting adapter: good portability |

## Portability is the key strategic question

When a provider changes its offering, can you move?

- **Closed-model fine-tunes** (Gemini, GPT, Nova) live only on that provider. You keep your **data and evaluation**, not the weights.
- **Open-model fine-tunes** on platforms that let you export adapters, or on your own GPUs, give you portable weights.

Mitigation for any path: keep your curated datasets, labeling guides, graders and evaluation harness **provider-neutral** (plain JSONL plus conversion scripts). Then re-tuning elsewhere is days, not months.

## Data formats differ

The same example in three shapes:

```jsonl
{"messages": [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "card blocked"}, {"role": "assistant", "content": "card_blocked"}]}
{"systemInstruction": {"role": "system", "parts": [{"text": "Classify the intent."}]}, "contents": [{"role": "user", "parts": [{"text": "card blocked"}]}, {"role": "model", "parts": [{"text": "card_blocked"}]}]}
{"prompt": [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "card blocked"}], "completion": [{"role": "assistant", "content": "card_blocked"}]}
```

The first is the OpenAI-style chat format also used by Azure; the second is the Gemini (Vertex AI) format; the third is TRL's prompt-completion format. Keep one canonical format and write converters.

## Hands-on: launching jobs (illustrative, check current SDK docs)

**Vertex AI (Google Gen AI SDK):**

```python
# pip install google-genai ; authenticate with Application Default Credentials; set
# GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION and GOOGLE_GENAI_USE_VERTEXAI=True
from google import genai
from google.genai.types import CreateTuningJobConfig, TuningDataset, TuningValidationDataset

client = genai.Client()
job = client.tunings.tune(
    base_model="gemini-2.5-flash",
    training_dataset=TuningDataset(gcs_uri="gs://your-bucket/intents/train.jsonl"),
    config=CreateTuningJobConfig(tuned_model_display_name="intents-v1",
                                 validation_dataset=TuningValidationDataset(gcs_uri="gs://your-bucket/intents/val.jsonl")),
)
print(job.name, job.state)   # poll client.tunings.get(name=job.name) until it finishes
```

**Amazon Bedrock (boto3):**

```python
import os, boto3
bedrock = boto3.client("bedrock", region_name=os.environ["AWS_REGION"])
resp = bedrock.create_model_customization_job(
    jobName="intents-sft-v1",
    customModelName="intents-v1",
    roleArn=os.environ["BEDROCK_CUSTOMIZATION_ROLE_ARN"],
    baseModelIdentifier=os.environ["BASE_MODEL_ID"],        # a customizable model in your region
    customizationType="FINE_TUNING",
    trainingDataConfig={"s3Uri": "s3://your-bucket/intents/train.jsonl"},
    validationDataConfig={"validators": [{"s3Uri": "s3://your-bucket/intents/val.jsonl"}]},
    outputDataConfig={"s3Uri": "s3://your-bucket/intents/output/"},
    hyperParameters={"epochCount": "2"},                    # names and ranges differ per base model
)
print(resp["jobArn"])
```

Required data formats and hyperparameter names differ per base model on Bedrock; follow the model-specific page.

## Decision checklist

1. Which base models does the provider let you tune, in which **regions**?
2. Which **methods** (SFT, preference, RFT, distillation)?
3. **Data handling:** is training data retained, used for other purposes, encrypted, region-bound?
4. **Inference:** price per token for custom models, any dedicated capacity requirement, latency.
5. **Lifecycle:** what happens when the base model is deprecated?
6. **Portability:** can you export weights or adapters?

## Worked example: choosing a path in Riyadh

A Saudi e-commerce firm needs an Arabic intent classifier with data staying in-kingdom. They check which providers offer tuning and inference in an in-kingdom region for their candidates, find their preferred closed model is not tunable there, and choose to fine-tune an open-weight model with LoRA on in-kingdom cloud GPUs, keeping adapters portable. Their canonical JSONL and evaluation harness let them compare against a hosted option later if regional availability changes.

## Pitfalls

- Building your product on one provider's fine-tuning without a migration plan.
- Ignoring region and data-retention terms.
- Forgetting that custom-model inference may be priced differently from base models.
- Assuming the next base-model version will accept your old fine-tune (it will not; retrain).

## How to measure success

A provider decision documented against the checklist, provider-neutral data and evaluation assets, and a tested conversion path to at least one alternative.

## Video lecture: Hosted fine-tuning: which platforms, which models, which trade-offs

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Hosted fine-tuning
2. Analogy: tailoring a suit
3. Snapshot (verify)
4. What changed
5. Portability
6. Formats
7. Launching jobs
8. Six-point checklist
9. Worked example: Riyadh e-commerce
10. Simple example: team on Google Cloud
11. Lifecycle planning
12. FAQ: is custom-model inference pricier?
13. Try this now
14. Watch me do it
15. Recap

## Lecture transcript

### Hosted fine-tuning

In twenty twenty-six, a lot of teams woke up to an email saying the fine-tuning platform they built on was being wound down. Their fine-tuned models would keep running for a while, but new training would end. The teams that shrugged were the ones who had kept their data and evaluations portable. In this lesson you will map the hosted fine-tuning landscape, the formats and APIs, and the checklist that keeps you able to move.

### Analogy: tailoring a suit

Here is an analogy. Hosted fine-tuning is like tailoring a suit at a department store. It is convenient, they have the machines, and the result fits well. But the suit stays in their store, and if the store closes the tailoring counter, you cannot take the pattern elsewhere. Self-managed fine-tuning on open weights is owning the pattern. Either way, keep your measurements, which are your data and your evaluation, in your own notebook.

### Snapshot (verify)

Here is the snapshot, as of September twenty twenty-six, and please verify it. Google's Vertex AI offers supervised fine-tuning for Gemini two point five Pro, Flash and Flash-Lite, plus preference tuning, and tuning for selected open models. Amazon Bedrock customizes Amazon Nova models, including supervised fine-tuning, distillation and reinforcement fine-tuning starting with Nova two Lite. Microsoft Foundry fine-tunes the GPT four point one and GPT four o families, with other methods and models on varying access.

### What changed

And the changes. OpenAI began winding down its self-serve fine-tuning platform this year. It closed to new organizations, is phasing out new jobs for existing customers, and keeps existing fine-tuned models available for inference until their base models are deprecated. Mistral's docs now place its self-serve fine-tuning API under legacy resources. Meanwhile, several inference platforms offer LoRA fine-tuning on open-weight models, often letting you download the adapter.

### Portability

That makes portability the key strategic question. A fine-tune of a closed model lives only with that provider. You keep your data and evaluation, not the weights. A fine-tune of an open model, on your GPUs or a platform that exports adapters, gives you portable weights. Either way, keep your datasets, labeling guides, graders and evaluation harness provider-neutral, and moving becomes days of work, not months.

### Formats

Formats differ. OpenAI-style chat format, also used by Azure, is a list of messages with roles system, user and assistant. Gemini on Vertex uses contents with parts, and the assistant role is called model, plus a separate system instruction. TRL uses prompt and completion. The lesson shows the same example in all three. Keep one canonical format and write small converters, so switching is a script, not a project.

### Launching jobs

The lesson includes two illustrative job launches. On Vertex AI, the Google Gen AI SDK's tunings dot tune call takes a base model, a training dataset in Cloud Storage and a config with a display name and validation data, and you poll until it finishes. On Bedrock, boto3's create model customization job takes a job name, role, base model, customization type, S3 locations and hyperparameters, whose names differ per base model. Check each SDK's current docs before running.

### Six-point checklist

Before choosing, run a six-point checklist. Which base models can you tune, and in which regions? Which methods? How is training data handled: retained, used for anything else, encrypted, region-bound? What does inference cost for custom models, and is dedicated capacity required? What happens when the base model is deprecated? And can you export weights or adapters?

### Worked example: Riyadh e-commerce

A worked example. A Saudi e-commerce firm needs an Arabic intent classifier with data staying in the kingdom. They check which providers offer tuning and inference in an in-kingdom region, find their preferred closed model is not tunable there, and choose LoRA on an open-weight model using in-kingdom cloud GPUs. Their adapters are portable, and their canonical data and harness mean they can compare a hosted option later if availability changes.

### Simple example: team on Google Cloud

A simple example. A marketing team already runs everything on Google Cloud and needs a model that writes product titles in a strict format. Vertex AI supports supervised tuning for the Gemini two point five models they already use. They convert their canonical data to the Gemini format with a small script, launch a job, and compare the tuned model with their best prompt. The tuned model wins on format accuracy. They keep the canonical data and converter in Git, so they could re-run the same experiment elsewhere.

### Lifecycle planning

One more consideration: lifecycle. Hosted base models are retired on a schedule, and a fine-tune of a retired base model stops working when the base does. When you pick a hosted path, note the base model's announced retirement policy, set a reminder well before any date, and budget for retraining on the successor, including re-running your full evaluation. With portable data and an automated evaluation, that migration becomes a routine task rather than an emergency.

### FAQ: is custom-model inference pricier?

A question worth asking every time: does a hosted fine-tune cost more to run than the base model? It can. Some platforms price inference on custom models differently, and some require dedicated capacity to serve them. Put those numbers into your cost model before training, not after. A fine-tune that saves prompt tokens but requires expensive dedicated hosting may not save money at your volume.

### Try this now

Try this now. Write a twenty-line converter from your canonical JSON lines to one other provider's format, for example from messages to Gemini's contents and parts. Run it on ten examples and validate them. Once you have one converter working, switching providers stops being scary.

### Watch me do it

Watch me do it. Our canonical data is prompt and completion JSON lines. I write a short converter to the Gemini format: the system message becomes the system instruction, user turns become contents with role user and a text part, and the assistant completion becomes role model. I convert ten examples and validate them against the format described in the docs. Then I upload the full files to a storage bucket in the approved region and launch the tuning job with the Google Gen AI SDK, passing the base model, training dataset and a validation dataset. While it runs, I fill in the six-point checklist: models and regions, methods, data handling from the terms page, inference pricing for tuned models, the deprecation policy for the base, and export options, which for this closed model are none. When the job finishes, I evaluate the tuned model with the same harness and the same test set as our open-model adapter, so the comparison is fair.

### Recap

Recap. Hosted fine-tuning is convenient and fast-moving. Know the current landscape, and verify it. Closed-model fine-tunes are not portable, so keep data, graders and evaluations provider-neutral. Check regions, data handling, inference pricing and lifecycle. Your next step: complete the checklist for two hosted options and one self-managed option, and write converters from your canonical format to each.

## Key takeaways

- Hosted fine-tuning changes fast; verify models, methods, regions and pricing
- Vertex AI tunes Gemini 2.5 models; Bedrock customizes Nova (incl. RFT); Foundry tunes GPT-4.1/4o families
- OpenAI began winding down its self-serve fine-tuning platform in 2026
- Closed-model fine-tunes are not portable; keep data and evals provider-neutral
- Check region, data handling, inference pricing and deprecation lifecycle

## Try it

Complete the six-point decision checklist for two hosted options and one self-managed option for your task, and write converters from your canonical JSONL to each format.

- [Previous: Evaluation before and after fine-tuning](https://optimizeall.com/learn/fine-tuning-and-custom-models/evaluation-before-and-after)
- [Next: Fine-tuning embedding models for better retrieval](https://optimizeall.com/learn/fine-tuning-and-custom-models/embeddings-fine-tuning)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
