---
title: "Messages, roles and conversation state across Claude…"
description: "The shared mental model Every modern LLM API takes a list of conversation turns plus instructions , and returns a response containing content blocks…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/messages-and-conversation-formats
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Core API calls across providers · lesson 3 of 19 · 16 min

# Messages, roles and conversation state across Claude, OpenAI and Gemini

## The shared mental model

Every modern LLM API takes a **list of conversation turns** plus **instructions**, and returns a **response containing content blocks**. The differences are naming and where state lives.

| Concept | Claude Messages API | OpenAI Responses API | Gemini `generate_content` |
|---|---|---|---|
| Call | `client.messages.create(...)` | `client.responses.create(...)` | `client.models.generate_content(...)` |
| Instructions | `system=` (top-level) | `instructions=` (or a developer/system message) | `config=GenerateContentConfig(system_instruction=...)` |
| Turns | `messages=[{"role": "user"/"assistant", "content": ...}]` | `input=` string or list of items (`role`, `content`) | `contents=` string, list, or `types.Content(role="user"/"model", parts=[...])` |
| Required length cap | `max_tokens` (required) | optional `max_output_tokens` | optional `max_output_tokens` in config |
| Output | `resp.content` list of blocks (`text`, `tool_use`, `thinking`...) | `resp.output` items; `resp.output_text` convenience | `resp.candidates[...]`; `resp.text` convenience |
| Why it stopped | `resp.stop_reason` | `resp.status` / item types | `candidates[0].finish_reason` |
| Usage | `resp.usage.input_tokens/output_tokens` | `resp.usage` | `resp.usage_metadata` |
| Server-side state | Stateless: you resend history | Optional: `previous_response_id` chains stored responses (`store` defaults to true) | Stateless calls; `client.chats` helper keeps history client-side |

OpenAI's older **Chat Completions API** (`client.chat.completions.create(messages=[...])`) remains supported and is what many "OpenAI-compatible" providers implement; the Responses API is OpenAI's recommended interface for new work.

## Roles and instruction hierarchy

- **System / developer instructions** set behavior, persona, constraints and format. Keep them stable (it helps caching) and put per-request data in the user turn.
- **User** turns carry the request and data.
- **Assistant / model** turns are the model's previous outputs. On current Claude models you cannot "prefill" the start of an assistant reply; use instructions or structured outputs instead.

## Managing conversation state yourself

For chat products, you store the transcript and resend it each turn (or chain with provider state). Decide:

- **What to keep**: all turns, a rolling window, or a summary plus recent turns.
- **Where**: your database keyed by conversation ID, with retention matching your privacy policy.
- **How much**: enforce a token budget per request; count tokens with each provider's token-counting endpoint for precision.

Storing state yourself makes provider switching easy and keeps you in control of retention; provider-side state (like `previous_response_id`) reduces payload size but ties the conversation to that provider and its retention settings.

## Worked example: one support turn, three providers

A UK subscription-box company answers "Can I pause my box for a month?" using its policy text.

```python
import os
import anthropic
from openai import OpenAI
from google import genai
from google.genai import types

SYSTEM = "You are a support assistant for BoxJoy UK. Answer only from the policy. If unsure, offer a human handover."
POLICY = "Pausing: customers can pause for up to 2 consecutive months from Account > Subscription."
QUESTION = "Can I pause my box for a month?"
user_turn = f"Policy:\n{POLICY}\n\nCustomer question: {QUESTION}"

# Claude
c = anthropic.Anthropic().messages.create(
    model=os.environ.get("CLAUDE_MODEL", "claude-sonnet-5"), max_tokens=500, system=SYSTEM,
    messages=[{"role": "user", "content": user_turn}])
print("Claude:", "".join(b.text for b in c.content if b.type == "text"), "| stop:", c.stop_reason)

# OpenAI (Responses API)
o = OpenAI().responses.create(
    model=os.environ.get("OPENAI_MODEL", "gpt-5.5"), instructions=SYSTEM, input=user_turn,
    store=False)                                    # don't keep this response server-side
print("OpenAI:", o.output_text)

# Gemini
g = genai.Client().models.generate_content(
    model=os.environ.get("GEMINI_MODEL", "gemini-flash-latest"), contents=user_turn,
    config=types.GenerateContentConfig(system_instruction=SYSTEM, max_output_tokens=500))
print("Gemini:", g.text)
```

## A normalized internal format

Define your own small message type and convert at the edges; module 5 builds this into a full adapter.

```python
from dataclasses import dataclass

@dataclass
class Turn:
    role: str      # "user" | "assistant"
    text: str

def to_claude(turns):  return [{"role": t.role, "content": t.text} for t in turns]
def to_openai(turns):  return [{"role": t.role, "content": t.text} for t in turns]
def to_gemini(turns):  return [types.Content(role="model" if t.role == "assistant" else "user",
                                             parts=[types.Part.from_text(text=t.text)]) for t in turns]
```

## Multi-turn example

For a follow-up ("And can I skip just one delivery?"), append the previous assistant answer and the new user message, then call again. Keep the policy in a stable system prompt or first user turn so caching can reuse it.

## Token counting before you send

For long conversations, count tokens before each call so you can trim proactively instead of hitting context or cost limits. Each provider offers a counting method: Claude's `client.messages.count_tokens(...)` takes the same model, system and messages as a real request; the Gemini SDK has `client.models.count_tokens(...)`; OpenAI documents input-token counting for the Responses API. Tokenizers differ by provider and even by model generation, so never reuse one provider's count for another, and re-baseline when you change models.

## Pitfalls

- Putting per-request data in the system prompt (breaks caching and mixes concerns).
- Forgetting `max_tokens` on Claude (it is required).
- Relying on provider-side conversation storage without checking retention and privacy terms.
- Assuming role names match across providers (`assistant` vs `model`).

## Measuring success

Tokens per turn, conversation length distribution, truncation or `max_tokens` stop rates, and the percentage of conversations you can replay against another provider (a good test of portability).

## Video lecture: Messages, roles and conversation state across Claude, OpenAI and Gemini

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Messages across providers
2. Why it matters
3. The shared model
4. Field by field
5. Two OpenAI APIs + a Claude note
6. Where state lives
7. History costs tokens
8. Simple example: gym reminder
9. Business example: pause-my-box
10. Normalize at the edges
11. Common mistakes
12. Stable first, variable last
13. Deeper: BoxJoy's provider test (illustrative)
14. Watch me do it: one question, three providers
15. Recap + try this now

## Lecture transcript

### Messages across providers

Three providers, three different ways of saying the same thing. One calls it system, another calls it instructions, the third calls it system instruction. One says assistant, another says model. If you've ever copied code between providers and watched it fail on a single field name, this lesson is for you. You'll learn the shared mental model behind every chat API, and exactly how Claude, OpenAI and Gemini express it.

### Why it matters

Why does this matter? Because the message format is the foundation of every feature you'll build: streaming, tools, structured outputs, caching. If you understand it once, every provider's docs become quick to read. Think of it like driving in different countries. The steering wheel might be on the left or right, and road signs look different, but a car is still a car. Learn the car, then learn each country's signs.

### The shared model

Here's the shared model. Every modern API takes instructions, which set behavior, plus a list of conversation turns, and returns a response made of content blocks. Instructions say who the assistant is and what rules it follows. User turns carry the request and data. Assistant turns are the model's earlier replies. The response tells you what was generated, why it stopped, and how many tokens were used. Hold that picture, and the differences become details.

### Field by field

Now the details. With Claude, you call messages create, pass system at the top level, a messages list with user and assistant roles, and a required max tokens value. The reply has a content list of blocks and a stop reason. With OpenAI's Responses API, you call responses create, pass instructions and an input, and read output text for convenience. With Gemini, you call models generate content with contents, and put the system instruction and max output tokens in a config object, then read the text property.

### Two OpenAI APIs + a Claude note

A word on OpenAI's two APIs. The Responses API is OpenAI's recommended interface for new work. The older Chat Completions API is still supported, and it matters beyond OpenAI because many other providers copy its shape as an OpenAI compatible endpoint. So you'll meet Chat Completions often, even when you're not calling OpenAI at all. Also note: on current Claude models you can't prefill the start of the assistant's reply, so use instructions or structured outputs to control format instead.

### Where state lives

Where does conversation state live? Claude's API is stateless: you resend the history every turn. Gemini calls are stateless too, with a chat helper that keeps history on your side. OpenAI's Responses API can store responses and chain them with a previous response id, and storage is on by default unless you set store to false. Storing history yourself gives you portability and control over retention. Provider side state reduces payload size, but ties the conversation to that provider and its retention settings.

### History costs tokens

Let's talk about tokens, because conversation state and cost are tightly linked. Every turn you resend is billed again as input. A support chat that runs to thirty turns can easily send ten times more tokens on its last turn than its first. So decide early how much history each request really needs. Many teams keep the last six to ten turns in full, plus a short summary of anything older, plus the pinned facts that must never be forgotten, like the customer's plan and country. Count tokens with each provider's counting endpoint, because tokenizers differ between providers and even between model generations.

### Simple example: gym reminder

A simple example. You ask all three providers, write a friendly one line reminder that the gym closes early on Friday. With Claude, you pass the request in the messages list and read the text block. With OpenAI, you pass it as input and read output text. With Gemini, you pass it as contents and read text. Same request, three shapes, three nearly identical answers. Once you've done this once, switching feels routine.

### Business example: pause-my-box

Now a realistic business example. A UK subscription box company answers, can I pause my box for a month, using its policy text. The system instructions say: answer only from the policy, and offer a human handover if unsure. The policy and question go in the user turn. The code in the lesson sends this to all three providers, sets store to false on OpenAI so the response isn't kept server side, and prints each answer. The team compares tone, accuracy and token usage before choosing.

### Normalize at the edges

To keep your code portable, define your own small message type, a role and some text, and convert it at the edges to each provider's format. Claude and OpenAI both accept user and assistant roles; Gemini needs model instead of assistant and wraps text in parts. Module five turns this into a full adapter layer. For multi turn chat, append the last answer and the new user message, and keep stable content like the policy in the same place each turn so caching can reuse it.

### Common mistakes

Common mistakes. Putting per request data in the system prompt, which breaks caching and mixes concerns. Forgetting max tokens on Claude. Relying on provider side storage without checking retention and privacy. And assuming role names match everywhere. Each one causes a confusing bug, usually at the worst possible moment, like during a demo.

### Stable first, variable last

One more practical tip. Decide deliberately where your stable content lives. The system instructions and any reference text that doesn't change between turns should sit at the start of the request, in the same position and with identical wording every time. Variable things, like today's date, the user's latest question and retrieved snippets, belong at the end. That single habit makes prompt caching work across providers, reduces cost and speeds up responses, and it costs you nothing but a little discipline.

### Deeper: BoxJoy's provider test (illustrative)

Let's deepen the BoxJoy UK example with illustrative numbers. The company handles about four thousand support chats a month. When they compared providers on two hundred real questions, all three answered the pause question correctly, but differed on edge cases, like pausing during a promotion. Because the policy lived in the user turn and the instructions in the system prompt, switching providers for the test took minutes. They also set store to false on OpenAI and kept history in their own database for thirty days, matching their privacy notice. The winning provider was chosen on edge case accuracy and cost per resolved chat, not on the easy questions everyone got right.

### Watch me do it: one question, three providers

Watch me do it. Let's walk through the three provider support example. The system string says: answer only from the policy, and offer a human if unsure. The policy and the question are combined into one user turn. For Claude, I pass system at the top level, the user turn in messages, and max tokens five hundred, then print the text and the stop reason. For OpenAI, instructions carries the system text, input carries the user turn, and store false keeps the response off OpenAI's servers; I print output text. For Gemini, contents is the user turn, and the config holds the system instruction and max output tokens; I print dot text. Then the normalized Turn class: role and text. To Claude and to OpenAI map it almost directly, while to Gemini renames assistant to model and wraps text in parts. Now a follow up question just means appending two turns and converting again.

### Recap + try this now

Quick recap. Every API is instructions plus turns in, content blocks out. Claude uses system, messages and a required max tokens. OpenAI's Responses API uses instructions and input. Gemini uses contents and a config. Own your conversation state for portability. Try this now: run the three provider example with your own policy text and question, then send a three turn conversation to two providers using the normalized format, and compare token usage.

## Key takeaways

- All three APIs take instructions plus turns and return content blocks; names and state handling differ.
- Claude uses system and messages with a required max_tokens; OpenAI Responses uses instructions and input; Gemini uses contents and config.
- OpenAI's Responses API is recommended for new work; Chat Completions remains widely supported and copied.
- Store conversation state yourself for portability and privacy control; provider state trades control for convenience.
- Normalize messages internally and convert at the edges.

## Try it

Send the same three-turn conversation to two providers using the normalized Turn format, and compare tokens used and answers.

- [Previous: API keys, authentication and secrets management](https://optimizeall.com/learn/ai-platform-apis-integration/api-keys-auth-and-secrets)
- [Next: Streaming responses to users and between services](https://optimizeall.com/learn/ai-platform-apis-integration/streaming-responses)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
