Voice AI & Conversational AgentsPlatforms and stacks · Lesson 5 of 17

OpenAI Realtime API, Gemini Live and open-source stacks

Article · 16 min · 9 min lecture

Video lecture

OpenAI Realtime API, Gemini Live and open-source stacks

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Code-first voice stacks

  • OpenAI Realtime + Agents SDK
  • Gemini Live API
  • LiveKit Agents, Pipecat
  • Orchestration platforms

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

When to go beyond an integrated platform

Integrated platforms are the fastest path, but you may want code-level control: a custom UI, your own orchestration, speech-to-speech models, special compliance hosting, or a product where voice is one part of a larger app. This lesson covers the main code-first options.

OpenAI Realtime API

The Realtime API provides low-latency, speech-to-speech interaction with OpenAI's realtime models (the gpt-realtime family; OpenAI released updated versions through 2025 and 2026, so check the models page for the current default and mini variants).

Connection options:

TransportUse when
WebRTCBrowser or mobile clients; handles microphone, echo cancellation and jitter for you
WebSocketServer-side apps where your server owns the audio stream (for example bridging telephony media streams)
SIPConnecting phone calls directly to the Realtime API via a SIP trunk

Key features: built-in turn detection (server_vad based on silence, or semantic_vad that judges whether the user has finished speaking), function calling (tools), remote MCP server support, image input, transcription of input audio, voices chosen from OpenAI's set, and session-level instructions.

Security pattern: never ship your standard API key to the browser. Your server creates a short-lived ephemeral client secret and the browser uses it to connect.

Hands-on: a browser voice agent with the OpenAI Agents SDK (TypeScript)

npm install @openai/agents zod

Server (Node) endpoint that mints an ephemeral key:

// server.ts (e.g. an Express route); OPENAI_API_KEY stays on the server
app.post("/realtime-token", async (_req, res) => {
  const r = await fetch("https://api.openai.com/v1/realtime/client_secrets", {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}`, "Content-Type": "application/json" },
    body: JSON.stringify({ session: { type: "realtime", model: process.env.REALTIME_MODEL ?? "gpt-realtime" } }),
  });
  if (!r.ok) return res.status(502).json({ error: "token_failed" });
  const data = await r.json();
  res.json({ value: data.value });
});

Browser client:

import { RealtimeAgent, RealtimeSession, tool } from "@openai/agents/realtime";
import { z } from "zod";

const checkAvailability = tool({
  name: "check_availability",
  description: "Check free appointment slots for a date (YYYY-MM-DD).",
  parameters: z.object({ date: z.string() }),
  async execute({ date }) {
    const r = await fetch(`/api/availability?date=${encodeURIComponent(date)}`);
    if (!r.ok) return "Availability service is unavailable. Offer to take a callback number.";
    return await r.text();
  },
});

const agent = new RealtimeAgent({
  name: "Receptionist",
  instructions: "You are an AI receptionist for Nova Dental. Say you are an AI assistant in your first sentence. Keep replies short. Never give clinical advice.",
  tools: [checkAvailability],
});

const { value } = await (await fetch("/realtime-token", { method: "POST" })).json();
const session = new RealtimeSession(agent); // WebRTC in the browser: mic and speaker wired automatically
await session.connect({ apiKey: value });

The SDK's RealtimeSession handles audio, turn detection, interruptions, tool calls and handoffs between agents. API shapes evolve; follow the current Agents SDK voice quickstart and pin package versions.

Google Gemini Live API

Google's Gemini Live API offers low-latency bidirectional audio (and video) streaming with Gemini models, including native-audio models that generate speech directly, with function calling and voice activity detection. It is available through Google AI Studio and Vertex AI. Check the current model names, supported languages and regions in Google's documentation.

Open-source frameworks

  • LiveKit Agents: an open-source framework on top of LiveKit's WebRTC infrastructure, with plugins for many STT, LLM, TTS and realtime providers (including ElevenLabs, OpenAI, Google, Deepgram), a turn-detection model and telephony via SIP.
  • Pipecat: an open-source Python framework for voice and multimodal pipelines, composable "frame processors", many provider integrations, and transports such as WebRTC and telephony WebSockets.

They suit teams that need full control, self-hosting, custom turn logic or unusual integrations. You take on more engineering and operations.

Orchestration platforms

Platforms such as Vapi and Retell let you mix providers (for example Deepgram STT, your LLM, ElevenLabs TTS) behind an API with telephony, tools and analytics. They sit between integrated platforms and raw frameworks.

Decision matrix

NeedBest fit
Launch in days, non-developer teamIntegrated platform (e.g. ElevenAgents)
Speech-to-speech in your own web appOpenAI Realtime via Agents SDK, or Gemini Live
Mix-and-match providers with telephony, moderate codeOrchestration platform
Full control, self-hosting, custom pipelinesLiveKit Agents or Pipecat
Specific brand voice with speech-to-speech feelCascaded stack with fast components, or a hybrid

Worked example: a UK language-tutoring startup

A startup wants students to practice spoken English with an AI tutor that corrects gently and handles interruptions naturally. They choose speech-to-speech (natural prosody matters more than a custom voice), build in the browser with the Agents SDK over WebRTC, mint ephemeral tokens server-side, and add a tool that logs vocabulary to the student's profile. For safeguarding, they add session limits for under-18 users and store transcripts only with consent.

Pitfalls

  • Exposing a standard API key in client code.
  • Assuming realtime models behave like text models; test tool-calling accuracy and instruction-following in audio.
  • Ignoring audio-token pricing; long sessions with verbose agents can cost more than expected. Check current pricing.

Key takeaways

  • The OpenAI Realtime API offers speech-to-speech via WebRTC, WebSocket or SIP, with server_vad or semantic_vad turn detection, tools and remote MCP.
  • Browser clients must use short-lived ephemeral client secrets minted by your server, never the standard API key.
  • Gemini Live API, LiveKit Agents, Pipecat and orchestration platforms (Vapi, Retell) cover other points on the control-versus-speed spectrum.
  • Choose by team skills, need for custom voice, control, hosting and cost; test tool-calling accuracy in audio.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A web app connects to the OpenAI Realtime API from the browser. How should it authenticate?
  2. Which transport suits a server that bridges raw phone audio streams to a realtime model?
  3. A team needs self-hosting and fully custom turn logic. Which option fits best?

Put it into practice

Run the Agents SDK example locally with an ephemeral token endpoint. Ask the agent to book a slot in three different phrasings and check whether it calls the tool with correct parameters each time.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.