---
title: "OpenAI Realtime API, Gemini Live and open-source stacks"
description: "When to go beyond an integrated platform Integrated platforms are the fastest path, but you may want code-level control: a custom UI, your own…"
url: https://optimizeall.com/learn/voice-ai-agents/openai-realtime-and-other-stacks
updated: 2026-10-05
---

Voice AI & Conversational Agents · Platforms and stacks · lesson 5 of 17 · 16 min

# OpenAI Realtime API, Gemini Live and open-source stacks

## When to go beyond an integrated platform

Integrated platforms are the fastest path, but you may want code-level control: a custom UI, your own orchestration, speech-to-speech models, special compliance hosting, or a product where voice is one part of a larger app. This lesson covers the main code-first options.

## OpenAI Realtime API

The Realtime API provides low-latency, speech-to-speech interaction with OpenAI's realtime models (the **gpt-realtime** family; OpenAI released updated versions through 2025 and 2026, so check the models page for the current default and mini variants).

**Connection options:**

| Transport | Use when |
|---|---|
| **WebRTC** | Browser or mobile clients; handles microphone, echo cancellation and jitter for you |
| **WebSocket** | Server-side apps where your server owns the audio stream (for example bridging telephony media streams) |
| **SIP** | Connecting phone calls directly to the Realtime API via a SIP trunk |

**Key features:** built-in turn detection (`server_vad` based on silence, or `semantic_vad` that judges whether the user has finished speaking), function calling (tools), remote MCP server support, image input, transcription of input audio, voices chosen from OpenAI's set, and session-level instructions.

**Security pattern:** never ship your standard API key to the browser. Your server creates a short-lived **ephemeral client secret** and the browser uses it to connect.

## Hands-on: a browser voice agent with the OpenAI Agents SDK (TypeScript)

```bash
npm install @openai/agents zod
```

Server (Node) endpoint that mints an ephemeral key:

```typescript
// server.ts (e.g. an Express route); OPENAI_API_KEY stays on the server
app.post("/realtime-token", async (_req, res) => {
  const r = await fetch("https://api.openai.com/v1/realtime/client_secrets", {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.OPENAI_API_KEY}`, "Content-Type": "application/json" },
    body: JSON.stringify({ session: { type: "realtime", model: process.env.REALTIME_MODEL ?? "gpt-realtime" } }),
  });
  if (!r.ok) return res.status(502).json({ error: "token_failed" });
  const data = await r.json();
  res.json({ value: data.value });
});
```

Browser client:

```typescript
import { RealtimeAgent, RealtimeSession, tool } from "@openai/agents/realtime";
import { z } from "zod";

const checkAvailability = tool({
  name: "check_availability",
  description: "Check free appointment slots for a date (YYYY-MM-DD).",
  parameters: z.object({ date: z.string() }),
  async execute({ date }) {
    const r = await fetch(`/api/availability?date=${encodeURIComponent(date)}`);
    if (!r.ok) return "Availability service is unavailable. Offer to take a callback number.";
    return await r.text();
  },
});

const agent = new RealtimeAgent({
  name: "Receptionist",
  instructions: "You are an AI receptionist for Nova Dental. Say you are an AI assistant in your first sentence. Keep replies short. Never give clinical advice.",
  tools: [checkAvailability],
});

const { value } = await (await fetch("/realtime-token", { method: "POST" })).json();
const session = new RealtimeSession(agent); // WebRTC in the browser: mic and speaker wired automatically
await session.connect({ apiKey: value });
```

The SDK's `RealtimeSession` handles audio, turn detection, interruptions, tool calls and handoffs between agents. API shapes evolve; follow the current Agents SDK voice quickstart and pin package versions.

## Google Gemini Live API

Google's **Gemini Live API** offers low-latency bidirectional audio (and video) streaming with Gemini models, including native-audio models that generate speech directly, with function calling and voice activity detection. It is available through Google AI Studio and Vertex AI. Check the current model names, supported languages and regions in Google's documentation.

## Open-source frameworks

- **LiveKit Agents**: an open-source framework on top of LiveKit's WebRTC infrastructure, with plugins for many STT, LLM, TTS and realtime providers (including ElevenLabs, OpenAI, Google, Deepgram), a turn-detection model and telephony via SIP.
- **Pipecat**: an open-source Python framework for voice and multimodal pipelines, composable "frame processors", many provider integrations, and transports such as WebRTC and telephony WebSockets.

They suit teams that need full control, self-hosting, custom turn logic or unusual integrations. You take on more engineering and operations.

## Orchestration platforms

Platforms such as **Vapi** and **Retell** let you mix providers (for example Deepgram STT, your LLM, ElevenLabs TTS) behind an API with telephony, tools and analytics. They sit between integrated platforms and raw frameworks.

## Decision matrix

| Need | Best fit |
|---|---|
| Launch in days, non-developer team | Integrated platform (e.g. ElevenAgents) |
| Speech-to-speech in your own web app | OpenAI Realtime via Agents SDK, or Gemini Live |
| Mix-and-match providers with telephony, moderate code | Orchestration platform |
| Full control, self-hosting, custom pipelines | LiveKit Agents or Pipecat |
| Specific brand voice with speech-to-speech feel | Cascaded stack with fast components, or a hybrid |

## Worked example: a UK language-tutoring startup

A startup wants students to practice spoken English with an AI tutor that corrects gently and handles interruptions naturally. They choose speech-to-speech (natural prosody matters more than a custom voice), build in the browser with the Agents SDK over WebRTC, mint ephemeral tokens server-side, and add a tool that logs vocabulary to the student's profile. For safeguarding, they add session limits for under-18 users and store transcripts only with consent.

## Pitfalls

- Exposing a standard API key in client code.
- Assuming realtime models behave like text models; test tool-calling accuracy and instruction-following in audio.
- Ignoring audio-token pricing; long sessions with verbose agents can cost more than expected. Check current pricing.

## Video lecture: OpenAI Realtime API, Gemini Live and open-source stacks

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

1. Code-first voice stacks
2. Why code-first options
3. OpenAI Realtime API
4. Analogy: opening a restaurant
5. Features and security
6. Agents SDK in the browser
7. Gemini Live API
8. Open source and orchestration
9. Worked example: tutoring startup
10. Example 2: Karachi call-center outsourcer
11. Common mistakes
12. Watch me do it: browser realtime agent
13. Recap and next step
14. Try this now

## Lecture transcript

### Code-first voice stacks

Integrated platforms get you live fast. But sometimes you need code-level control: your own interface, your own orchestration, speech to speech models, special hosting for compliance, or voice as one feature in a bigger app. In this lesson, you'll learn the OpenAI Realtime API and Agents SDK, Google's Gemini Live API, the open-source frameworks LiveKit Agents and Pipecat, and orchestration platforms, and you'll finish with a decision matrix for choosing between them.

### Why code-first options

Why learn code-first options if platforms exist? Because some products need them: voice inside your own app, special hosting for compliance, speech to speech natural flow, or unusual integrations. And even if you use a platform today, knowing the alternatives helps you negotiate, plan for growth, and avoid being locked into a single vendor's roadmap.

### OpenAI Realtime API

OpenAI's Realtime API gives low-latency speech to speech with the G P T realtime model family, which OpenAI has updated through twenty twenty-five and twenty twenty-six, so check which version is current. You can connect three ways. WebRTC for browsers and mobile, where microphone, echo cancellation and jitter are handled for you. WebSocket for servers that own the audio stream, for example when bridging phone audio. And SIP, to connect phone calls straight to the API.

### Analogy: opening a restaurant

An analogy for these options: building a voice product is like opening a restaurant. An integrated platform is a franchise: menu, kitchen and systems provided, open in weeks. An orchestration platform is a shared commercial kitchen: you bring your own recipes and suppliers, they provide the equipment. A realtime model API is hiring a star chef who does everything in one flow. And open-source frameworks are building your own kitchen from scratch: total control, but you fix the plumbing too.

### Features and security

Key features include built-in turn detection, with a server V A D mode based on silence and a semantic V A D mode that judges whether the user has actually finished. It supports function calling, remote MCP servers, image input and transcription of what the user says. And one security rule matters most: never ship your standard API key to a browser. Your server creates a short-lived ephemeral client secret, and the browser uses that to connect.

### Agents SDK in the browser

The OpenAI Agents SDK makes this compact. On the server, a small route calls the client secrets endpoint and returns the ephemeral value. In the browser, you define tools with a name, description, a Zod schema for parameters, and an execute function. You create a realtime agent with instructions, including saying it's an AI assistant in its first sentence, and your tools. Then you create a realtime session and connect with the ephemeral key. The session wires up the microphone and speaker, turn detection, interruptions, tool calls and handoffs. The full code is in the lesson text.

### Gemini Live API

Google's Gemini Live API offers low-latency, two-way audio and video streaming with Gemini models, including native audio models that generate speech directly, plus function calling and voice activity detection. It's available through Google AI Studio and Vertex AI. As always, check current model names, languages and regions. It's a strong option if you're already in the Google Cloud ecosystem or need video input alongside voice.

### Open source and orchestration

If you want full control, look at open source. LiveKit Agents runs on LiveKit's WebRTC infrastructure, with plugins for many speech, model and voice providers, including ElevenLabs, OpenAI, Google and Deepgram, a turn detection model, and telephony through SIP. Pipecat is an open-source Python framework that composes voice pipelines from processors, with many integrations and transports. Between those and integrated platforms sit orchestration platforms like Vapi and Retell, which let you mix providers behind one API, with telephony and analytics included.

### Worked example: tutoring startup

Here's a UK language-tutoring startup. Students practice spoken English with an AI tutor that corrects gently and handles interruptions naturally. Natural rhythm matters more than a custom voice, so they choose speech to speech, built in the browser with the Agents SDK over WebRTC, with ephemeral tokens minted on their server. A tool logs new vocabulary to each student's profile. For safeguarding, they add session limits for under eighteens and only store transcripts with consent.

### Example 2: Karachi call-center outsourcer

A second example. A Karachi call-center outsourcer wants to run voice agents for several clients, with each client's data kept separate and some audio processed on its own servers. They choose an open-source framework with self-hosted speech recognition for Urdu, and plug in different language models and voices per client. It takes an engineering team weeks, not days, but gives them control over data, costs and customization that a single platform couldn't. For them, the extra effort is the product.

### Common mistakes

Common mistakes with code-first stacks. Exposing a standard API key in browser code, which anyone can copy. Assuming realtime models behave like text models, when tool-calling accuracy and instruction-following should be tested in audio, with real accents. Ignoring audio-token pricing on long sessions with chatty agents. And choosing full control when you don't have the engineers to maintain it, which turns a quick win into a long project.

### Watch me do it: browser realtime agent

Watch me do it. I build a browser version of Aria with the OpenAI Agents SDK, to compare with our platform agent. In a small project I install the agents package and Zod. On the server, I write the token route from the lesson text: it posts to the client secrets endpoint with our real key from an environment variable and returns only the short-lived value. I test it with curl and see a value come back. In the browser file, I define the check availability tool: a name, a description, a Zod object with a date string, and an execute function that calls our mock availability endpoint and returns a friendly error message if it fails. I create a realtime agent with instructions that begin: you are an AI receptionist for Nova Dental, say you are an AI assistant in your first sentence, keep replies short, never give clinical advice. I create the session, fetch the token, and connect. The browser asks for the microphone. I say: hi, can I get a cleaning next Tuesday? The agent introduces itself as an AI assistant, calls the tool with next Tuesday's date, and offers two times. I try two more phrasings, tomorrow afternoon and the fifteenth, and log whether each tool call had the right date. Two of three were right; the fifteenth picked the wrong month, so I add today's date to the instructions and retest.

### Recap and next step

Recap and decision time. Non-developer team, launch in days: an integrated platform. Speech to speech in your own web app: OpenAI Realtime with the Agents SDK, or Gemini Live. Mix and match providers with telephony and moderate code: an orchestration platform. Full control and self-hosting: LiveKit Agents or Pipecat. Watch out for exposed API keys, untested tool-calling in audio, and audio-token costs on long sessions. Your next step: run the Agents SDK example locally, then ask it to book a slot three different ways and see if it calls the tool correctly.

### Try this now

Try this now, step by step. First, install the Agents SDK in a small web project. Second, add a server route that mints an ephemeral client secret, keeping your real key in an environment variable. Third, create a realtime agent with short instructions that start with an AI disclosure. Fourth, add one tool with a Zod schema that calls a mock availability endpoint. Fifth, connect in the browser, then ask for a booking three different ways and note whether the tool is called with the right date each time.

## Key takeaways

- The OpenAI Realtime API offers speech-to-speech via WebRTC, WebSocket or SIP, with server_vad or semantic_vad turn detection, tools and remote MCP.
- Browser clients must use short-lived ephemeral client secrets minted by your server, never the standard API key.
- Gemini Live API, LiveKit Agents, Pipecat and orchestration platforms (Vapi, Retell) cover other points on the control-versus-speed spectrum.
- Choose by team skills, need for custom voice, control, hosting and cost; test tool-calling accuracy in audio.

## Try it

Run the Agents SDK example locally with an ephemeral token endpoint. Ask the agent to book a slot in three different phrasings and check whether it calls the tool with correct parameters each time.

- [Previous: Building an agent with ElevenAgents](https://optimizeall.com/learn/voice-ai-agents/building-with-elevenagents)
- [Next: Telephony: phone numbers, SIP trunks and Twilio](https://optimizeall.com/learn/voice-ai-agents/telephony-sip-and-twilio)
- [All lessons of Voice AI & Conversational Agents](https://optimizeall.com/learn/voice-ai-agents)
