Multimodal & Reasoning Models in PracticeAudio: speech-to-text and text-to-speech · Lesson 6 of 17

Text-to-speech and voice pipelines

Article · 11 min · 7 min lecture

Video lecture

Text-to-speech and voice pipelines

11 chapters · about 7 min · full transcript

Coming soon

Chapter 1 of 11

Text-to-speech and voice pipelines

  • Writing for the ear
  • Steerable voices
  • Pipelines and latency
  • Testing and consent

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Text-to-speech today

Text-to-speech (TTS) turns text into spoken audio. Modern neural TTS can sound remarkably natural, with control over voice, pace, emotion and pronunciation, and many services support multiple languages. (Our separate course on AI video and voice covers specific creative tools; this lesson focuses on the underlying concepts, pipelines and evaluation.)

Writing for the ear

Text written for reading often sounds wrong when spoken. Before blaming the voice, fix the script:

  • Short sentences. Long clauses are hard to follow by ear.
  • Spell out tricky items: "twenty twenty-six" or "2026" depending on how the engine reads it; test numbers, dates, currencies and abbreviations.
  • Remove visual formatting: bullet symbols, tables and URLs read awkwardly.
  • Signal structure verbally: "There are three steps. First..."
  • Pronunciation control: many services support phonetic hints, custom pronunciation dictionaries or markup such as SSML (Speech Synthesis Markup Language) for pauses, emphasis and pronunciation. Support varies by provider.

A useful prompt when generating scripts with a language model:

Write this as a 45-second voiceover script. Spoken style, sentences under
18 words, no lists or symbols, numbers written as they should be spoken,
and include [pause] where a natural breath would go.

Voice agents: the speech-to-speech loop

A voice assistant combines components:

User speaks -> STT (streaming) -> language model (+ tools) -> TTS (streaming) -> user hears

Some newer systems use models that process audio more directly (speech-to-speech), which can reduce latency and preserve tone, but the component view remains useful for understanding trade-offs.

Latency is the defining challenge. In conversation, people notice pauses quickly. Techniques to reduce perceived latency:

  • Streaming at every stage (partial transcripts, streamed model output, streamed audio).
  • Short first responses ("Let me check that for you") while tools run.
  • Faster models for conversational turns; heavier reasoning only when needed.
  • Good turn-taking: detecting when the user has finished speaking (endpointing) and handling interruptions ("barge-in") gracefully.

Evaluating TTS and voice agents

TTS quality is partly subjective. Practical evaluation:

  • Listening tests: have several people rate naturalness and clarity on a small scale, comparing voices blind where possible.
  • Pronunciation checklist: your brand names, product names, local place names, numbers and dates.
  • Intelligibility: transcribe the TTS output with an STT system and compare to the script; large differences flag unclear speech.
  • Voice agent metrics: response latency (median and slow tail), task completion, interruption handling, escalation rate and user satisfaction.

Worked example: appointment reminders

A clinic group in the UAE wants automated appointment reminder calls in Arabic and English.

  • Scripts are short, confirm date and time twice, and offer "press 1 to confirm, 2 to reschedule".
  • A pronunciation list covers clinic names and doctor names.
  • Dates are formatted to be spoken naturally in each language.
  • The system identifies itself as an automated assistant at the start.
  • Pilot listening tests with staff and a small group of patients lead to a slower speaking rate for older patients.

Voice technology carries specific risks:

  • Voice cloning must only be done with the explicit, informed consent of the person whose voice it is, and within the provider's policies.
  • Disclosure: people should know when they are talking to an automated system; some jurisdictions and platforms require it.
  • Impersonation and fraud: synthetic voices have been used in scams. Organisations should verify unusual voice requests (for example payment instructions) through a separate channel.
  • Accessibility: TTS can widen access for people with visual impairments or reading difficulties; design with them in mind.

Steerable voices

Current TTS systems increasingly accept natural-language direction as well as markup: "warm and reassuring, slightly slower, smile in the voice" or "energetic sports-commentary style". Some let you design a new voice from a description, and many offer expressive tags for pauses, laughter or whispering. Features and syntax differ by provider (ElevenLabs, OpenAI, Google and others), so keep a short style guide per voice and test it on your actual scripts.

Hands-on: generate a voiceover and check it with a round trip

# Illustrative pipeline: script -> TTS -> STT -> diff. Swap in your providers' SDK calls.
import difflib, os
from openai import OpenAI

client = OpenAI()
SCRIPT = ("Welcome to Bloom Café. From the first of October, we open at seven a.m. "
          "on weekdays. Pre-order on our app and skip the queue.")

speech = client.audio.speech.create(
    model=os.environ["TTS_MODEL"],            # a current TTS model from your provider's docs
    voice=os.environ.get("TTS_VOICE", "alloy"),
    input=SCRIPT,
    instructions="Warm, friendly, unhurried; a small smile in the voice.",
)
with open("bloom.mp3", "wb") as f:
    f.write(speech.content)

with open("bloom.mp3", "rb") as audio:
    heard = client.audio.transcriptions.create(model=os.environ["STT_MODEL"], file=audio).text

for line in difflib.unified_diff(SCRIPT.split(), heard.split(), lineterm="", n=0):
    print(line)

Words that come back different from the script (brand names, times, prices) are candidates for a pronunciation dictionary entry or a script rewrite ("seven a.m." versus "7am"). Check the current parameter names in your provider's audio documentation; voice lists and model names change.

Rules around synthetic voices keep tightening. Consumer protection and telemarketing rules in several jurisdictions restrict AI-generated voices in unsolicited calls, platforms require labels for realistic synthetic media, and the EU AI Act introduces transparency obligations for AI-generated audio and deepfakes that apply in stages. For any voice product: disclose automation at the start, keep consent records for any cloned voice, and take legal advice for outbound calling.

Going further

For production voice agents, log each stage's timing separately (STT final, model first token, TTS first audio byte) so you know where latency comes from. Test with real network conditions, including mobile connections, and with diverse speakers and accents. Accuracy and latency across user groups should be monitored, not assumed.

Key takeaways

  • Write for the ear: short sentences, spoken-form numbers, no visual formatting, verbal structure.
  • Voice agents chain STT, a language model and TTS (or use speech-to-speech models); latency and turn-taking are the core challenges.
  • Evaluate with listening tests, pronunciation checklists, STT round-trip intelligibility and conversation metrics.
  • Clone voices only with explicit consent, disclose automation, and guard against voice fraud.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which script change most improves TTS output?
  2. What is 'barge-in' in voice agents?
  3. What is required before cloning someone's voice?

Put it into practice

Write a 45-second voiceover script following the 'write for the ear' rules. Generate it with a TTS tool and run it back through STT: where do the transcript and script differ?

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.