Fine-Tuning, Distillation and Custom ModelsEvaluation, hosted platforms, embeddings and cost · Lesson 12 of 16

Hosted fine-tuning: which platforms, which models, which trade-offs

Article · 15 min · 8 min lecture

Video lecture

Hosted fine-tuning: which platforms, which models, which trade-offs

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Hosted fine-tuning

  • The 2026 landscape
  • Formats and APIs
  • Staying portable

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

A fast-moving market

Hosted fine-tuning lets you upload data, run a job and call a custom model without managing GPUs. It is also one of the fastest-changing parts of the AI market. The snapshot below reflects official documentation as of September 2026; verify current models, methods, regions and pricing before you commit, and design so you can move.

The landscape (September 2026 snapshot)

ProviderWhat is offered (verify)Notes
Google Cloud Vertex AI (Google's docs now also brand it the Gemini Enterprise Agent Platform)Supervised fine-tuning for Gemini 2.5 Pro, Gemini 2.5 Flash and Gemini 2.5 Flash-Lite (text, image, audio, video and document data per docs); preference tuning for Gemini; supervised and distillation tuning for selected open modelsVia the Google Gen AI SDK or REST; regional availability matters for residency
Amazon BedrockModel customization including supervised fine-tuning, distillation and reinforcement fine-tuning for Amazon Nova models (RFT starting with Nova 2 Lite); other base models vary by regionGraders for RFT via AWS Lambda or model-as-judge; check whether custom models need dedicated capacity for inference
Microsoft Foundry (Azure)Fine-tuning for GPT-4.1, GPT-4.1-mini, GPT-4.1-nano, GPT-4o and GPT-4o-mini with supervised methods, plus other methods and models with varying (sometimes gated) accessMicrosoft documentation noted GPT-5-family models were not supported for fine-tuning at the time of writing
OpenAISelf-serve fine-tuning (SFT, DPO, RFT, vision) historicallyIn 2026 OpenAI began winding down its self-serve fine-tuning platform: closed to new organizations, with job creation for existing customers being phased out; existing fine-tuned models remain available for inference until their base models are deprecated
Mistral AIHas offered a fine-tuning API for its modelsIts docs now place the self-serve fine-tuning API under legacy resources; check current status
Inference platforms for open modelsSeveral providers offer LoRA/SFT (and sometimes DPO) on open-weight models with serving includedYou can often download the resulting adapter: good portability

Portability is the key strategic question

When a provider changes its offering, can you move?

  • Closed-model fine-tunes (Gemini, GPT, Nova) live only on that provider. You keep your data and evaluation, not the weights.
  • Open-model fine-tunes on platforms that let you export adapters, or on your own GPUs, give you portable weights.

Mitigation for any path: keep your curated datasets, labeling guides, graders and evaluation harness provider-neutral (plain JSONL plus conversion scripts). Then re-tuning elsewhere is days, not months.

Data formats differ

The same example in three shapes:

{"messages": [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "card blocked"}, {"role": "assistant", "content": "card_blocked"}]}
{"systemInstruction": {"role": "system", "parts": [{"text": "Classify the intent."}]}, "contents": [{"role": "user", "parts": [{"text": "card blocked"}]}, {"role": "model", "parts": [{"text": "card_blocked"}]}]}
{"prompt": [{"role": "system", "content": "Classify the intent."}, {"role": "user", "content": "card blocked"}], "completion": [{"role": "assistant", "content": "card_blocked"}]}

The first is the OpenAI-style chat format also used by Azure; the second is the Gemini (Vertex AI) format; the third is TRL's prompt-completion format. Keep one canonical format and write converters.

Hands-on: launching jobs (illustrative, check current SDK docs)

Vertex AI (Google Gen AI SDK):

# pip install google-genai ; authenticate with Application Default Credentials; set
# GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION and GOOGLE_GENAI_USE_VERTEXAI=True
from google import genai
from google.genai.types import CreateTuningJobConfig, TuningDataset, TuningValidationDataset

client = genai.Client()
job = client.tunings.tune(
    base_model="gemini-2.5-flash",
    training_dataset=TuningDataset(gcs_uri="gs://your-bucket/intents/train.jsonl"),
    config=CreateTuningJobConfig(tuned_model_display_name="intents-v1",
                                 validation_dataset=TuningValidationDataset(gcs_uri="gs://your-bucket/intents/val.jsonl")),
)
print(job.name, job.state)   # poll client.tunings.get(name=job.name) until it finishes

Amazon Bedrock (boto3):

import os, boto3
bedrock = boto3.client("bedrock", region_name=os.environ["AWS_REGION"])
resp = bedrock.create_model_customization_job(
    jobName="intents-sft-v1",
    customModelName="intents-v1",
    roleArn=os.environ["BEDROCK_CUSTOMIZATION_ROLE_ARN"],
    baseModelIdentifier=os.environ["BASE_MODEL_ID"],        # a customizable model in your region
    customizationType="FINE_TUNING",
    trainingDataConfig={"s3Uri": "s3://your-bucket/intents/train.jsonl"},
    validationDataConfig={"validators": [{"s3Uri": "s3://your-bucket/intents/val.jsonl"}]},
    outputDataConfig={"s3Uri": "s3://your-bucket/intents/output/"},
    hyperParameters={"epochCount": "2"},                    # names and ranges differ per base model
)
print(resp["jobArn"])

Required data formats and hyperparameter names differ per base model on Bedrock; follow the model-specific page.

Decision checklist

  1. Which base models does the provider let you tune, in which regions?
  2. Which methods (SFT, preference, RFT, distillation)?
  3. Data handling: is training data retained, used for other purposes, encrypted, region-bound?
  4. Inference: price per token for custom models, any dedicated capacity requirement, latency.
  5. Lifecycle: what happens when the base model is deprecated?
  6. Portability: can you export weights or adapters?

Worked example: choosing a path in Riyadh

A Saudi e-commerce firm needs an Arabic intent classifier with data staying in-kingdom. They check which providers offer tuning and inference in an in-kingdom region for their candidates, find their preferred closed model is not tunable there, and choose to fine-tune an open-weight model with LoRA on in-kingdom cloud GPUs, keeping adapters portable. Their canonical JSONL and evaluation harness let them compare against a hosted option later if regional availability changes.

Pitfalls

  • Building your product on one provider's fine-tuning without a migration plan.
  • Ignoring region and data-retention terms.
  • Forgetting that custom-model inference may be priced differently from base models.
  • Assuming the next base-model version will accept your old fine-tune (it will not; retrain).

How to measure success

A provider decision documented against the checklist, provider-neutral data and evaluation assets, and a tested conversion path to at least one alternative.

Key takeaways

  • Hosted fine-tuning changes fast; verify models, methods, regions and pricing
  • Vertex AI tunes Gemini 2.5 models; Bedrock customizes Nova (incl. RFT); Foundry tunes GPT-4.1/4o families
  • OpenAI began winding down its self-serve fine-tuning platform in 2026
  • Closed-model fine-tunes are not portable; keep data and evals provider-neutral
  • Check region, data handling, inference pricing and deprecation lifecycle

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A team fine-tunes a closed model on a hosted platform. If the provider ends the offering, what do they keep?
  2. Which statement about OpenAI fine-tuning in 2026 is accurate?
  3. Why keep a single canonical training format with converters?

Put it into practice

Complete the six-point decision checklist for two hosted options and one self-managed option for your task, and write converters from your canonical JSONL to each format.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.