Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APICore API calls across providers · Lesson 4 of 19

Streaming responses to users and between services

Article · 15 min · 9 min lecture

Video lecture

Streaming responses to users and between services

16 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 16

Streaming responses

  • Why it matters
  • Provider stream shapes
  • Relaying to the browser
  • Production gotchas

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why stream

A response that takes 12 seconds to complete feels broken if the user stares at a spinner, but fine if words start appearing in under a second. Streaming improves perceived latency (time to first token), lets you show progress for long tasks, and avoids HTTP timeouts on long outputs. Most providers deliver streams as Server-Sent Events (SSE).

Provider streaming shapes

Claude (Python): use the messages.stream helper, which accumulates the final message for you.

import anthropic
client = anthropic.Anthropic()
with client.messages.stream(model="claude-sonnet-5", max_tokens=2000,
                            messages=[{"role": "user", "content": "Draft a 150-word LinkedIn post about our new Riyadh office."}]) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()          # full message with usage and stop_reason
print("\n", final.usage.output_tokens, final.stop_reason)

Under the hood, Claude emits events such as message_start, content_block_start, content_block_delta (with text_delta, thinking_delta or input_json_delta for tool arguments), content_block_stop, message_delta (with the stop reason and usage) and message_stop.

Claude (TypeScript):

import Anthropic from '@anthropic-ai/sdk';
const client = new Anthropic();
const stream = client.messages.stream({
  model: 'claude-sonnet-5', max_tokens: 2000,
  messages: [{ role: 'user', content: 'Draft a 150-word LinkedIn post about our new Riyadh office.' }],
});
for await (const event of stream) {
  if (event.type === 'content_block_delta' && event.delta.type === 'text_delta') process.stdout.write(event.delta.text);
}
const final = await stream.finalMessage();

OpenAI (Responses API): pass stream=True and handle typed events; text arrives as response.output_text.delta events.

from openai import OpenAI
client = OpenAI()
stream = client.responses.create(model="gpt-5.5", input="Draft a 150-word LinkedIn post...", stream=True)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print("\nusage:", event.response.usage)

Gemini: generate_content_stream yields chunks with a .text property.

from google import genai
client = genai.Client()
for chunk in client.models.generate_content_stream(model="gemini-flash-latest",
                                                   contents="Draft a 150-word LinkedIn post..."):
    print(chunk.text or "", end="", flush=True)

Always confirm model IDs and event names against current docs; event taxonomies grow as features are added.

Relaying streams to a browser

Your backend holds the API key, so the stream flows provider → your server → browser. A minimal FastAPI relay using SSE:

# pip install fastapi uvicorn anthropic
import json
import anthropic
from fastapi import FastAPI
from fastapi.responses import StreamingResponse

app = FastAPI()
client = anthropic.AsyncAnthropic()

@app.get("/api/draft")
async def draft(topic: str):
    async def events():
        try:
            async with client.messages.stream(model="claude-sonnet-5", max_tokens=1500,
                    messages=[{"role": "user", "content": f"Write a short post about: {topic}"}]) as s:
                async for text in s.text_stream:
                    yield f"data: {json.dumps({'delta': text})}\n\n"
                final = await s.get_final_message()
                yield f"data: {json.dumps({'done': True, 'output_tokens': final.usage.output_tokens})}\n\n"
        except anthropic.APIError as exc:
            yield f"data: {json.dumps({'error': 'generation_failed'})}\n\n"   # don't leak internals
    return StreamingResponse(events(), media_type="text/event-stream",
                             headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})

In the browser, new EventSource('/api/draft?topic=...') receives each data: line. Add authentication and per-user rate limits before exposing it.

Streaming gotchas

  • Proxies and buffering: some reverse proxies buffer responses; disable buffering for the stream route (the X-Accel-Buffering: no header helps with Nginx).
  • Errors mid-stream: a stream can fail after sending partial text; show the partial output with a clear retry option, and log the error.
  • Usage accounting: token usage usually arrives at the end of the stream; record it from the final event.
  • Tool calls while streaming: tool arguments arrive as partial JSON deltas; only parse and execute once the block is complete (or use SDK helpers that accumulate them).
  • Moderation and validation: if you must validate output (for example structured JSON or policy checks), either validate at the end before committing actions, or stream to the user with a visible "draft" state.

Worked example: a KSA real-estate listing writer

Agents in Riyadh generate Arabic and English listing descriptions. Before streaming, users waited 8–10 seconds (illustrative) and often clicked twice, doubling cost. After streaming with a relay, first words appear quickly, the button disables while streaming, and duplicate requests dropped. Usage from the final event feeds per-agent cost reports.

A frontend checklist for streamed text

  • Disable the submit button while a stream is active, and show a stop button that aborts the request.
  • Render Markdown incrementally only if your renderer tolerates incomplete syntax; otherwise render plain text until the stream ends.
  • Keep the partial output if the stream fails, and label it clearly as incomplete.
  • For right-to-left languages such as Arabic and Urdu, set the text direction on the output container so streamed text flows correctly.

Pitfalls

  • Streaming directly from the browser to the provider (exposes keys).
  • Not handling client disconnects, so generation continues and bills after the user leaves; cancel upstream when the client goes away.
  • Parsing partial JSON as if complete.

Measuring success

Time to first token (p50/p95), abandonment rate during generation, duplicate-submit rate, and stream error rate.

Key takeaways

  • Streaming improves perceived latency and avoids timeouts on long outputs.
  • Claude: messages.stream with text_stream and get_final_message; OpenAI: stream=True with typed events; Gemini: generate_content_stream.
  • Relay streams through your backend as SSE so keys stay private.
  • Handle proxy buffering, mid-stream errors, client disconnects and end-of-stream usage.
  • Only parse tool arguments or JSON once the block is complete.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What does streaming primarily improve for users?
  2. In OpenAI's Responses API stream, which event carries incremental text?
  3. A user closes the tab mid-generation. What should your relay do?

Put it into practice

Build the FastAPI SSE relay, connect a simple web page with EventSource, and log time to first token and output tokens for ten requests.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.