---
title: "Streaming responses to users and between services"
description: "Why stream A response that takes 12 seconds to complete feels broken if the user stares at a spinner, but fine if words start appearing in under a…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/streaming-responses
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Core API calls across providers · lesson 4 of 19 · 15 min

# Streaming responses to users and between services

## Why stream

A response that takes 12 seconds to complete feels broken if the user stares at a spinner, but fine if words start appearing in under a second. Streaming improves **perceived latency** (time to first token), lets you show progress for long tasks, and avoids HTTP timeouts on long outputs. Most providers deliver streams as **Server-Sent Events (SSE)**.

## Provider streaming shapes

**Claude (Python)**: use the `messages.stream` helper, which accumulates the final message for you.

```python
import anthropic
client = anthropic.Anthropic()
with client.messages.stream(model="claude-sonnet-5", max_tokens=2000,
                            messages=[{"role": "user", "content": "Draft a 150-word LinkedIn post about our new Riyadh office."}]) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()          # full message with usage and stop_reason
print("\n", final.usage.output_tokens, final.stop_reason)
```

Under the hood, Claude emits events such as `message_start`, `content_block_start`, `content_block_delta` (with `text_delta`, `thinking_delta` or `input_json_delta` for tool arguments), `content_block_stop`, `message_delta` (with the stop reason and usage) and `message_stop`.

**Claude (TypeScript)**:

```typescript
import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const stream = client.messages.stream({
  model: 'claude-sonnet-5', max_tokens: 2000,
  messages: [{ role: 'user', content: 'Draft a 150-word LinkedIn post about our new Riyadh office.' }],
});
for await (const event of stream) {
  if (event.type === 'content_block_delta' && event.delta.type === 'text_delta') process.stdout.write(event.delta.text);
}
const final = await stream.finalMessage();
```

**OpenAI (Responses API)**: pass `stream=True` and handle typed events; text arrives as `response.output_text.delta` events.

```python
from openai import OpenAI
client = OpenAI()
stream = client.responses.create(model="gpt-5.5", input="Draft a 150-word LinkedIn post...", stream=True)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print("\nusage:", event.response.usage)
```

**Gemini**: `generate_content_stream` yields chunks with a `.text` property.

```python
from google import genai
client = genai.Client()
for chunk in client.models.generate_content_stream(model="gemini-flash-latest",
                                                   contents="Draft a 150-word LinkedIn post..."):
    print(chunk.text or "", end="", flush=True)
```

Always confirm model IDs and event names against current docs; event taxonomies grow as features are added.

## Relaying streams to a browser

Your backend holds the API key, so the stream flows provider → your server → browser. A minimal FastAPI relay using SSE:

```python
# pip install fastapi uvicorn anthropic
import json
import anthropic
from fastapi import FastAPI
from fastapi.responses import StreamingResponse

app = FastAPI()
client = anthropic.AsyncAnthropic()

@app.get("/api/draft")
async def draft(topic: str):
    async def events():
        try:
            async with client.messages.stream(model="claude-sonnet-5", max_tokens=1500,
                    messages=[{"role": "user", "content": f"Write a short post about: {topic}"}]) as s:
                async for text in s.text_stream:
                    yield f"data: {json.dumps({'delta': text})}\n\n"
                final = await s.get_final_message()
                yield f"data: {json.dumps({'done': True, 'output_tokens': final.usage.output_tokens})}\n\n"
        except anthropic.APIError as exc:
            yield f"data: {json.dumps({'error': 'generation_failed'})}\n\n"   # don't leak internals
    return StreamingResponse(events(), media_type="text/event-stream",
                             headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})
```

In the browser, `new EventSource('/api/draft?topic=...')` receives each `data:` line. Add authentication and per-user rate limits before exposing it.

## Streaming gotchas

- **Proxies and buffering**: some reverse proxies buffer responses; disable buffering for the stream route (the `X-Accel-Buffering: no` header helps with Nginx).
- **Errors mid-stream**: a stream can fail after sending partial text; show the partial output with a clear retry option, and log the error.
- **Usage accounting**: token usage usually arrives at the end of the stream; record it from the final event.
- **Tool calls while streaming**: tool arguments arrive as partial JSON deltas; only parse and execute once the block is complete (or use SDK helpers that accumulate them).
- **Moderation and validation**: if you must validate output (for example structured JSON or policy checks), either validate at the end before committing actions, or stream to the user with a visible "draft" state.

## Worked example: a KSA real-estate listing writer

Agents in Riyadh generate Arabic and English listing descriptions. Before streaming, users waited 8–10 seconds (illustrative) and often clicked twice, doubling cost. After streaming with a relay, first words appear quickly, the button disables while streaming, and duplicate requests dropped. Usage from the final event feeds per-agent cost reports.

## A frontend checklist for streamed text

- Disable the submit button while a stream is active, and show a stop button that aborts the request.
- Render Markdown incrementally only if your renderer tolerates incomplete syntax; otherwise render plain text until the stream ends.
- Keep the partial output if the stream fails, and label it clearly as incomplete.
- For right-to-left languages such as Arabic and Urdu, set the text direction on the output container so streamed text flows correctly.

## Pitfalls

- Streaming directly from the browser to the provider (exposes keys).
- Not handling client disconnects, so generation continues and bills after the user leaves; cancel upstream when the client goes away.
- Parsing partial JSON as if complete.

## Measuring success

Time to first token (p50/p95), abandonment rate during generation, duplicate-submit rate, and stream error rate.

## Video lecture: Streaming responses to users and between services

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

1. Streaming responses
2. Why it matters
3. Progress streaming
4. How streaming works
5. Claude streaming
6. OpenAI and Gemini streaming
7. Simple example: CLI email drafter
8. Relay through your backend
9. Business example: Riyadh listings
10. Production gotchas
11. Common mistakes
12. Streaming structured data
13. Measure streaming quality
14. Deeper: Riyadh listings (illustrative)
15. Watch me do it: the SSE relay
16. Recap + try this now

## Lecture transcript

### Streaming responses

Picture two apps that both take ten seconds to write a reply. In the first, you stare at a spinner for ten seconds. In the second, words start appearing in under a second, and you're reading while it writes. Same total time, completely different experience. That's streaming. In this lesson you'll learn how each provider streams, how to relay streams safely to a browser, and the gotchas that trip up production apps.

### Why it matters

Why does this matter? Users judge speed by the first sign of life, not the finish line. Think of a restaurant that brings bread while your meal cooks. The kitchen isn't faster, but the wait feels shorter, and you're less likely to walk out. Streaming is the bread basket for AI apps. It also has a technical benefit: long outputs can exceed HTTP timeouts if you wait for the whole response, and streaming avoids that.

### Progress streaming

There's a second kind of streaming worth knowing about: progress streaming. For multi step tasks, like an agent that searches, reads and then writes, users don't just want words appearing. They want to know what's happening. Searching your CRM. Found three contacts. Drafting the email. Those status events can travel over the same connection as the text. Think of a parcel tracking page. The parcel isn't faster, but seeing each step builds trust and stops people refreshing the page, which, in AI apps, often means paying twice.

### How streaming works

How does it work? Most providers send server sent events: a long lived HTTP response where the server pushes small messages, one after another, each describing a piece of the output. The first events announce the message has started, then many small deltas carry text, then a final event carries the stop reason and token usage. Your code reads the events as they arrive and displays or forwards the text.

### Claude streaming

With Claude in Python, the stream helper is the easiest path. You open it in a with block, loop over text stream to print text as it arrives, and call get final message at the end for the full message, stop reason and usage. In TypeScript, you iterate the events and write each text delta, then await final message. Underneath, Claude sends message start, content block deltas for text, thinking or tool arguments, a message delta with the stop reason, and message stop.

### OpenAI and Gemini streaming

OpenAI's Responses API streams when you pass stream equals true. You loop over typed events, and text arrives in events called response output text delta, with the finished response and usage in the completed event. Gemini's approach is the simplest: call generate content stream and each chunk has a text property you can print. Event names grow as providers add features, so check the docs, and write your handler to ignore event types it doesn't recognize.

### Simple example: CLI email drafter

A simple example. You're building a command line helper that drafts emails. Without streaming, it prints nothing for eight seconds, and you wonder if it crashed. With streaming, the first words appear almost immediately, and you can even press Control C to stop a draft that's heading the wrong way, saving tokens. Five lines of code changed, and the tool feels twice as fast.

### Relay through your backend

In real products, the browser never talks to the provider directly, because the key would be exposed. Instead, your backend opens the provider stream and relays it to the browser, usually as server sent events. The lesson's FastAPI example does exactly that: it streams text deltas as small JSON data lines, sends a final line with output tokens, and on errors sends a generic error message without leaking internals. The browser listens with EventSource.

### Business example: Riyadh listings

Now a realistic business example, with illustrative numbers. Property agents in Riyadh generate Arabic and English listing descriptions. Before streaming, they waited eight to ten seconds and often clicked generate twice, doubling cost. After adding a streaming relay, the first words appear quickly, the button is disabled while streaming, and duplicate requests dropped sharply. Usage from the final event now feeds a cost report per agent, so managers can see who uses AI most.

### Production gotchas

Let's cover the production gotchas. Proxies sometimes buffer responses, so disable buffering on the stream route. Streams can fail halfway, so show the partial text with a clear retry button. Usage usually arrives at the end, so record it from the final event. Tool arguments stream as partial JSON, so only parse them once the block is complete. And if the user closes the tab, cancel the upstream request, or you'll keep paying for text nobody will read.

### Common mistakes

Common mistakes, briefly. Calling the provider directly from the browser and exposing your key. Ignoring client disconnects, so generation continues after the user leaves. And parsing partial JSON as if it were complete, which causes random crashes that are hard to reproduce.

### Streaming structured data

A question that comes up with streaming and structured data: can you stream JSON? You can, but the partial JSON isn't valid until the end, so don't parse it as you go unless your SDK provides a tolerant partial parser. A good pattern for user interfaces is to stream a human readable draft first, then send a final validated structured object at the end, which your app uses for saving or routing. Users see progress, and your code only acts on validated data.

### Measure streaming quality

How do you know streaming is working well? Measure time to first token, the gap between the request leaving your server and the first text arriving. Track it at the median and the ninety fifth percentile, per model and per region. Watch abandonment, meaning users who leave before the stream ends, and duplicate submits. If time to first token jumps after a change, check whether you've added slow work before the model call, like retrieval or a large prompt, and consider streaming a status event while that work runs.

### Deeper: Riyadh listings (illustrative)

Let's deepen the Riyadh listings example with illustrative numbers. Two hundred property agents generated about six thousand descriptions a month. Before streaming, about one generation in six was a duplicate because impatient agents clicked again. After the streaming relay and a disabled button, duplicates almost disappeared, which saved more money than any prompt optimization had. The team also noticed a quality benefit: agents started reading the Arabic draft as it streamed and pressing stop when it went off track, then adjusting the brief. The final usage event fed a per agent cost report, and managers used it to coach the heaviest users on writing better briefs.

### Watch me do it: the SSE relay

Watch me do it. Let's walk through the FastAPI relay. The endpoint takes a topic query parameter. Inside, an async generator opens the Claude stream with the async client. For each text chunk from text stream, it yields an SSE line: the word data, a colon, and a small JSON object with the delta, followed by a blank line. When the stream ends, it awaits the final message and yields one last line with done true and the output token count. If an API error happens, it yields a generic error object, so the browser never sees internal details. The endpoint returns a streaming response with the event stream media type and two headers: no cache, and X accel buffering no, so Nginx doesn't buffer it. In the browser, an EventSource listens, appends each delta to the page, and closes on done. I run it and watch the words arrive.

### Recap + try this now

Quick recap. Streaming improves perceived speed and avoids timeouts. Claude has a stream helper, OpenAI streams typed events, and Gemini streams chunks. Relay through your backend, and handle buffering, failures, usage and disconnects. Try this now: build the FastAPI relay from the lesson, connect a tiny web page using EventSource, and log time to first token and output tokens for ten requests. Compare it with a non streaming version and feel the difference.

## Key takeaways

- Streaming improves perceived latency and avoids timeouts on long outputs.
- Claude: messages.stream with text_stream and get_final_message; OpenAI: stream=True with typed events; Gemini: generate_content_stream.
- Relay streams through your backend as SSE so keys stay private.
- Handle proxy buffering, mid-stream errors, client disconnects and end-of-stream usage.
- Only parse tool arguments or JSON once the block is complete.

## Try it

Build the FastAPI SSE relay, connect a simple web page with EventSource, and log time to first token and output tokens for ten requests.

- [Previous: Messages, roles and conversation state across Claude, OpenAI and Gemini](https://optimizeall.com/learn/ai-platform-apis-integration/messages-and-conversation-formats)
- [Next: Tool calling across Claude, OpenAI and Gemini](https://optimizeall.com/learn/ai-platform-apis-integration/tool-calling-across-providers)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
