Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APICore API calls across providers · Lesson 3 of 19

Messages, roles and conversation state across Claude, OpenAI and Gemini

Article · 16 min · 9 min lecture

Video lecture

Messages, roles and conversation state across Claude, OpenAI and Gemini

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Messages across providers

  • The shared mental model
  • Field-by-field differences
  • Managing conversation state

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The shared mental model

Every modern LLM API takes a list of conversation turns plus instructions, and returns a response containing content blocks. The differences are naming and where state lives.

ConceptClaude Messages APIOpenAI Responses APIGemini generate_content
Callclient.messages.create(...)client.responses.create(...)client.models.generate_content(...)
Instructionssystem= (top-level)instructions= (or a developer/system message)config=GenerateContentConfig(system_instruction=...)
Turnsmessages=[{"role": "user"/"assistant", "content": ...}]input= string or list of items (role, content)contents= string, list, or types.Content(role="user"/"model", parts=[...])
Required length capmax_tokens (required)optional max_output_tokensoptional max_output_tokens in config
Outputresp.content list of blocks (text, tool_use, thinking...)resp.output items; resp.output_text convenienceresp.candidates[...]; resp.text convenience
Why it stoppedresp.stop_reasonresp.status / item typescandidates[0].finish_reason
Usageresp.usage.input_tokens/output_tokensresp.usageresp.usage_metadata
Server-side stateStateless: you resend historyOptional: previous_response_id chains stored responses (store defaults to true)Stateless calls; client.chats helper keeps history client-side

OpenAI's older Chat Completions API (client.chat.completions.create(messages=[...])) remains supported and is what many "OpenAI-compatible" providers implement; the Responses API is OpenAI's recommended interface for new work.

Roles and instruction hierarchy

  • System / developer instructions set behavior, persona, constraints and format. Keep them stable (it helps caching) and put per-request data in the user turn.
  • User turns carry the request and data.
  • Assistant / model turns are the model's previous outputs. On current Claude models you cannot "prefill" the start of an assistant reply; use instructions or structured outputs instead.

Managing conversation state yourself

For chat products, you store the transcript and resend it each turn (or chain with provider state). Decide:

  • What to keep: all turns, a rolling window, or a summary plus recent turns.
  • Where: your database keyed by conversation ID, with retention matching your privacy policy.
  • How much: enforce a token budget per request; count tokens with each provider's token-counting endpoint for precision.

Storing state yourself makes provider switching easy and keeps you in control of retention; provider-side state (like previous_response_id) reduces payload size but ties the conversation to that provider and its retention settings.

Worked example: one support turn, three providers

A UK subscription-box company answers "Can I pause my box for a month?" using its policy text.

import os
import anthropic
from openai import OpenAI
from google import genai
from google.genai import types

SYSTEM = "You are a support assistant for BoxJoy UK. Answer only from the policy. If unsure, offer a human handover."
POLICY = "Pausing: customers can pause for up to 2 consecutive months from Account > Subscription."
QUESTION = "Can I pause my box for a month?"
user_turn = f"Policy:\n{POLICY}\n\nCustomer question: {QUESTION}"

# Claude
c = anthropic.Anthropic().messages.create(
    model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=500, system=SYSTEM,
    messages=[{"role": "user", "content": user_turn}])
print("Claude:", "".join(b.text for b in c.content if b.type == "text"), "| stop:", c.stop_reason)

# OpenAI (Responses API)
o = OpenAI().responses.create(
    model=os.environ.get("OPENAI_MODEL", "gpt-5.5"), instructions=SYSTEM, input=user_turn,
    store=False)                                    # don't keep this response server-side
print("OpenAI:", o.output_text)

# Gemini
g = genai.Client().models.generate_content(
    model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"), contents=user_turn,
    config=types.GenerateContentConfig(system_instruction=SYSTEM, max_output_tokens=500))
print("Gemini:", g.text)

A normalized internal format

Define your own small message type and convert at the edges; module 5 builds this into a full adapter.

from dataclasses import dataclass

@dataclass
class Turn:
    role: str      # "user" | "assistant"
    text: str

def to_claude(turns):  return [{"role": t.role, "content": t.text} for t in turns]
def to_openai(turns):  return [{"role": t.role, "content": t.text} for t in turns]
def to_gemini(turns):  return [types.Content(role="model" if t.role == "assistant" else "user",
                                             parts=[types.Part.from_text(text=t.text)]) for t in turns]

Multi-turn example

For a follow-up ("And can I skip just one delivery?"), append the previous assistant answer and the new user message, then call again. Keep the policy in a stable system prompt or first user turn so caching can reuse it.

Token counting before you send

For long conversations, count tokens before each call so you can trim proactively instead of hitting context or cost limits. Each provider offers a counting method: Claude's client.messages.count_tokens(...) takes the same model, system and messages as a real request; the Gemini SDK has client.models.count_tokens(...); OpenAI documents input-token counting for the Responses API. Tokenizers differ by provider and even by model generation, so never reuse one provider's count for another, and re-baseline when you change models.

Pitfalls

  • Putting per-request data in the system prompt (breaks caching and mixes concerns).
  • Forgetting max_tokens on Claude (it is required).
  • Relying on provider-side conversation storage without checking retention and privacy terms.
  • Assuming role names match across providers (assistant vs model).

Measuring success

Tokens per turn, conversation length distribution, truncation or max_tokens stop rates, and the percentage of conversations you can replay against another provider (a good test of portability).

Key takeaways

  • All three APIs take instructions plus turns and return content blocks; names and state handling differ.
  • Claude uses system and messages with a required max_tokens; OpenAI Responses uses instructions and input; Gemini uses contents and config.
  • OpenAI's Responses API is recommended for new work; Chat Completions remains widely supported and copied.
  • Store conversation state yourself for portability and privacy control; provider state trades control for convenience.
  • Normalize messages internally and convert at the edges.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which parameter is required on every Claude Messages API call?
  2. In the Gemini API, what role name is used for the model's previous turns?
  3. What is a benefit of storing conversation history in your own database rather than provider-side?

Put it into practice

Send the same three-turn conversation to two providers using the normalized Turn format, and compare tokens used and answers.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.