Multimodal & Reasoning Models in PracticeAudio: speech-to-text and text-to-speech · Lesson 5 of 17

Speech-to-text: transcription that holds up

Article · 11 min · 7 min lecture

Video lecture

Speech-to-text: transcription that holds up

11 chapters · about 7 min · full transcript

Coming soon

Chapter 1 of 11

Speech-to-text that holds up

  • Capabilities to look for
  • Measuring accuracy properly
  • Failure modes and fixes
  • Consent and compliance

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What speech-to-text does

Speech-to-text (STT), also called automatic speech recognition (ASR), converts spoken audio into text. Modern systems handle many languages, accents and noisy recordings far better than older generations, and some multimodal models can take audio directly and answer questions about it. Transcription is the foundation of meeting notes, call analytics, voice assistants, subtitles and searchable media archives.

Key capabilities to look for

  • Language coverage and code-switching. Many real conversations mix languages (for example English with Urdu, Arabic or Hindi). Test whether the system handles switching mid-sentence.
  • Speaker diarisation. Labelling who spoke when ("Speaker 1", "Speaker 2"). Essential for meetings and calls.
  • Timestamps. Word- or segment-level times for subtitles and navigation.
  • Custom vocabulary. Boosting recognition of product names, people's names and jargon.
  • Streaming vs batch. Real-time transcription for live captions and voice agents; batch for recordings, often cheaper and more accurate.

Measuring accuracy: word error rate and beyond

The standard metric is word error rate (WER): the number of substitutions, deletions and insertions needed to turn the transcript into the reference, divided by the number of words in the reference.

WER = (Substitutions + Deletions + Insertions) / Words in reference

Reference:  "send the invoice to Amina by Friday"
Transcript: "send an invoice to Amna Friday"
Substitutions: the->an, Amina->Amna (2); Deletions: by (1); Insertions: 0
WER = 3 / 7 = about 43%

WER treats all words equally, but business impact does not. A misheard "fifteen" vs "fifty" in a sales call matters far more than a missed "um". Add task-specific checks: accuracy of names, numbers, dates, product terms and action items.

Common failure modes

  • Names and jargon: people's names, brand names and technical terms get "corrected" into common words.
  • Numbers: "fifteen/fifty", currency amounts, phone numbers.
  • Overlapping speech and cross-talk in meetings.
  • Accents and dialects under-represented in training data, leading to uneven accuracy across speakers. This is also a fairness issue: check accuracy per group of speakers you serve.
  • Audio quality: echo, background noise, phone-line compression.
  • Hallucinated text: some systems can produce fluent text during silence or noise that was never spoken. Check transcripts of silent or noisy segments.

Improving results

  1. Fix the audio first. Better microphones, headsets and quieter rooms often beat any model change.
  2. Use custom vocabulary or prompts where supported to supply names and terms.
  3. Post-process with a language model carefully: it can fix punctuation, format numbers and apply a glossary, but instruct it not to change meaning, and keep the raw transcript for audit.
  4. Human review for high-stakes uses (legal, medical, compliance).

Worked example: sales call analytics

A B2B agency transcribes discovery calls to extract pain points, budget and next steps. Early issues: client company names mangled, budgets misheard. Improvements:

  • A vocabulary list per call with the client's company and product names from the CRM.
  • A post-processing step that extracts numbers with timestamps, so a rep can click and listen to confirm.
  • Extraction prompts that quote the transcript for each field.
  • Consent: callers are informed that calls are recorded and transcribed, in line with local law and company policy.

Recording and transcribing people is regulated differently across countries and sectors. Common requirements include informing participants, obtaining consent in some jurisdictions, limiting retention and protecting sensitive content. Check the rules that apply where you and your participants are, and document your approach.

The 2026 speech-to-text landscape

Options now fall into three groups:

  • Dedicated transcription APIs from AI labs and speech specialists (for example OpenAI's transcription models, ElevenLabs Scribe, Deepgram, AssemblyAI and the cloud providers' speech services), typically with diarisation, timestamps and vocabulary features.
  • Multimodal models that accept audio directly (such as Gemini), which can transcribe and also answer questions, summarise or extract from the audio in one call.
  • Open-weight models you can self-host for privacy or cost reasons.

Model line-ups change quickly. For example, in August 2026 OpenAI announced that several of its transcription models (including whisper-1 and the gpt-4o transcribe family) will be removed from its API on 26 February 2027, with newer transcription models as the migration path. Always check deprecation pages and keep your pipeline provider-agnostic.

Hands-on: compute WER and entity accuracy

# pip install jiwer
import re, jiwer

def normalise(t: str) -> str:
    t = t.lower()
    t = re.sub(r"[^\w\s']", " ", t)          # drop punctuation
    return re.sub(r"\s+", " ", t).strip()

reference = "send the invoice to Amina by Friday for fifteen thousand rupees"
hypothesis = "send an invoice to Amna Friday for fifty thousand rupees"

wer = jiwer.wer(normalise(reference), normalise(hypothesis))
print(f"WER: {wer:.0%}")

ENTITIES = ["amina", "friday", "fifteen thousand"]
hyp = normalise(hypothesis)
missed = [e for e in ENTITIES if e not in hyp]
print("Entity accuracy:", f"{1 - len(missed) / len(ENTITIES):.0%}", "missed:", missed)

This small example shows why entity checks matter: "fifteen" versus "fifty" is one substitution in WER terms and a costly mistake in business terms.

Prompting audio-capable models

When a multimodal model takes audio directly, give it the same briefing you would give a human transcriber:

Transcribe this 20-minute sales call between our rep (Sara) and a client
(Mr Qureshi). Label speakers by name. The call mixes English and Urdu:
transcribe Urdu in Urdu script and add an English translation in brackets.
Keep numbers as digits. Mark inaudible sections as [inaudible mm:ss].
Vocabulary: Optimize All, Shopify, CRO, ROAS.

Then keep the raw transcript and treat any summary or extraction as a separate, checkable step.

Going further

Build an evaluation set of 20 to 50 audio clips reflecting your real conditions: accents, languages, noise, jargon. Create reference transcripts carefully (this is labour-intensive, but essential). Measure WER overall and per segment (speaker group, noise level), plus field-level accuracy for the entities you care about. Re-run whenever you change provider or settings.

Key takeaways

  • STT turns speech into text; look for language and code-switching support, diarisation, timestamps, custom vocabulary, streaming vs batch.
  • Word error rate is the standard metric, but also measure accuracy on names, numbers and key entities.
  • Common failures: names, numbers, overlapping speech, uneven accuracy across accents, noise and hallucinated text.
  • Fix audio first, supply vocabulary, post-process carefully, and respect consent and retention rules.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A transcript has 2 substitutions, 1 deletion and 1 insertion against a 20-word reference. What is the WER?
  2. Why is WER alone insufficient for sales call analytics?
  3. What is often the cheapest way to improve transcription accuracy?

Put it into practice

Record or collect five short clips representative of your use case, write reference transcripts, and compute WER and name/number accuracy for one STT tool.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.