Multimodal & Reasoning Models in PracticeAudio: speech-to-text and text-to-speech · Lesson 5 of 17
Speech-to-text: transcription that holds up
Video lecture
Speech-to-text: transcription that holds up
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Speech-to-text that holds up
Send the invoice to Amina by Friday for fifteen thousand rupees. A transcription system hears: send an invoice to Amna Friday for fifty thousand rupees. On paper that is only a few word errors. In business, it is a wrong name, a lost deadline and a tenfold pricing mistake. In this lecture you will learn what modern speech-to-text can do, the capabilities to look for, how to measure accuracy properly, the common failure modes, how to prompt audio-capable models, and the consent rules that apply when you record people.
0:39 The landscape
Speech-to-text, also called automatic speech recognition, turns spoken audio into text. It is the foundation of meeting notes, call analytics, voice assistants, subtitles and searchable archives. In twenty twenty-six, options fall into three groups. Dedicated transcription APIs from AI labs and speech specialists, typically with speaker labels, timestamps and vocabulary features. Multimodal models that accept audio directly, like Gemini, which can transcribe and also answer questions about the audio in one call. And open-weight models you can host yourself for privacy or cost.
1:15 Plan for change
Model line-ups change fast, and that affects your roadmap. For example, in August twenty twenty-six OpenAI told developers that several transcription models, including whisper one and the GPT four o transcribe family, will be removed from its API on the twenty-sixth of February twenty twenty-seven, with newer transcription models as the migration path. The lesson is not about one vendor. It is that every speech pipeline should be provider-agnostic, with a thin wrapper around the transcription call and an evaluation set ready, so switching is a measured decision rather than an emergency.
1:55 Capabilities to check
When comparing systems, look for five capabilities. Language coverage and code-switching, because many real conversations mix English with Urdu, Arabic or Hindi mid-sentence. Speaker diarisation, labelling who spoke when. Timestamps at word or segment level. Custom vocabulary, to boost product names, people's names and jargon. And streaming versus batch: real-time for live captions and voice agents, batch for recordings, which is often cheaper and more accurate.
2:24 Measuring accuracy
The standard metric is word error rate: substitutions, deletions and insertions needed to turn the transcript into the reference, divided by the number of reference words. In our opening example, that is three errors over seven words in a shorter version of the sentence, around forty-three percent. But word error rate treats all words equally, and business impact does not. So add entity checks: accuracy on names, numbers, dates, product terms and action items. The lesson's hands-on code uses the jiwer library to compute word error rate after normalising punctuation and case, then checks a list of key entities and reports which were missed.
3:09 Example 1: a voice note
A simple worked example. You record a voice note while walking: remind me to call Doctor Farooqui on Thursday about the MRI results, and pick up the prescription from Boots. A transcription reads: call doctor Faruki on Thursday about the MRI result, and pick up the prescription from boots. The meaning mostly survived, but the name is misspelled and the pharmacy name lost its capital, which matters if you search for it later. Add a short vocabulary list, Farooqui, Boots, and the transcript comes back right. For personal notes that is convenience. For customer records, it is accuracy.
3:52 Example 2: an Urdu call centre (illustrative)
Now a business scenario, with illustrative numbers. A property developer's call centre in Lahore handles about fifteen thousand calls a month, mostly in Urdu with English property terms. They want transcripts to extract the caller's budget, preferred area and plot size. On a test of fifty calls with careful reference transcripts, the first system's word error rate is around twenty-two percent, and budget amounts are correct in only seventy percent of calls, mostly lakh and crore figures misheard. They try two other systems and a multimodal model with a transcription brief that names the speakers, the language mix, and asks for amounts as digits with the unit. The best option reaches about ninety-two percent budget accuracy, even though its overall word error rate only improves modestly. They choose it on entity accuracy, not word error rate, and add click-to-listen timestamps for every extracted amount. Illustrative figures.
4:55 Failure modes
Common failure modes: names and jargon corrected into common words; numbers, like fifteen versus fifty; overlapping speech in meetings; uneven accuracy across accents and dialects, which is also a fairness issue, so check accuracy per group of speakers you serve; poor audio from echo, noise or phone compression; and hallucinated text, where some systems produce fluent words during silence or noise that were never spoken. Check transcripts of silent segments specifically.
5:26 Improving results
Improvements, in order of payoff. Fix the audio first: better microphones, headsets and quieter rooms often beat any model change. Supply vocabulary or prompts with names and terms where supported. Post-process with a language model carefully, to fix punctuation, format numbers and apply a glossary, while instructing it not to change meaning, and keep the raw transcript for audit. And use human review for legal, medical or compliance uses. When a multimodal model takes audio directly, brief it like a human transcriber: who the speakers are, the languages and scripts to use, how to format numbers, how to mark inaudible sections, and a vocabulary list.
6:12 Worked example: sales calls
A worked example: a B2B agency transcribes discovery calls to extract pain points, budgets and next steps. Early issues: client company names mangled and budgets misheard. Fixes: a vocabulary list per call from the CRM; a step that extracts numbers with timestamps, so a rep can click and listen to confirm; extraction prompts that quote the transcript for each field; and consent. Callers are informed that calls are recorded and transcribed, in line with local law and company policy. Recording rules differ by country and sector, so document your approach, limit retention and protect sensitive content.
6:54 Recap
To recap. Choose among dedicated APIs, audio-capable multimodal models and self-hosted models, and keep your pipeline provider-agnostic because line-ups change. Check languages, diarisation, timestamps, vocabulary and streaming. Measure word error rate plus entity accuracy, fix audio first, and respect consent. Try this now: collect five short clips that represent your use case, write reference transcripts, and compute word error rate and name and number accuracy for one tool. Next: text-to-speech and voice pipelines.
What speech-to-text does
Speech-to-text (STT), also called automatic speech recognition (ASR), converts spoken audio into text. Modern systems handle many languages, accents and noisy recordings far better than older generations, and some multimodal models can take audio directly and answer questions about it. Transcription is the foundation of meeting notes, call analytics, voice assistants, subtitles and searchable media archives.
Key capabilities to look for
- Language coverage and code-switching. Many real conversations mix languages (for example English with Urdu, Arabic or Hindi). Test whether the system handles switching mid-sentence.
- Speaker diarisation. Labelling who spoke when ("Speaker 1", "Speaker 2"). Essential for meetings and calls.
- Timestamps. Word- or segment-level times for subtitles and navigation.
- Custom vocabulary. Boosting recognition of product names, people's names and jargon.
- Streaming vs batch. Real-time transcription for live captions and voice agents; batch for recordings, often cheaper and more accurate.
Measuring accuracy: word error rate and beyond
The standard metric is word error rate (WER): the number of substitutions, deletions and insertions needed to turn the transcript into the reference, divided by the number of words in the reference.
WER = (Substitutions + Deletions + Insertions) / Words in reference
Reference: "send the invoice to Amina by Friday"
Transcript: "send an invoice to Amna Friday"
Substitutions: the->an, Amina->Amna (2); Deletions: by (1); Insertions: 0
WER = 3 / 7 = about 43%WER treats all words equally, but business impact does not. A misheard "fifteen" vs "fifty" in a sales call matters far more than a missed "um". Add task-specific checks: accuracy of names, numbers, dates, product terms and action items.
Common failure modes
- Names and jargon: people's names, brand names and technical terms get "corrected" into common words.
- Numbers: "fifteen/fifty", currency amounts, phone numbers.
- Overlapping speech and cross-talk in meetings.
- Accents and dialects under-represented in training data, leading to uneven accuracy across speakers. This is also a fairness issue: check accuracy per group of speakers you serve.
- Audio quality: echo, background noise, phone-line compression.
- Hallucinated text: some systems can produce fluent text during silence or noise that was never spoken. Check transcripts of silent or noisy segments.
Improving results
- Fix the audio first. Better microphones, headsets and quieter rooms often beat any model change.
- Use custom vocabulary or prompts where supported to supply names and terms.
- Post-process with a language model carefully: it can fix punctuation, format numbers and apply a glossary, but instruct it not to change meaning, and keep the raw transcript for audit.
- Human review for high-stakes uses (legal, medical, compliance).
Worked example: sales call analytics
A B2B agency transcribes discovery calls to extract pain points, budget and next steps. Early issues: client company names mangled, budgets misheard. Improvements:
- A vocabulary list per call with the client's company and product names from the CRM.
- A post-processing step that extracts numbers with timestamps, so a rep can click and listen to confirm.
- Extraction prompts that quote the transcript for each field.
- Consent: callers are informed that calls are recorded and transcribed, in line with local law and company policy.
Consent and compliance
Recording and transcribing people is regulated differently across countries and sectors. Common requirements include informing participants, obtaining consent in some jurisdictions, limiting retention and protecting sensitive content. Check the rules that apply where you and your participants are, and document your approach.
The 2026 speech-to-text landscape
Options now fall into three groups:
- Dedicated transcription APIs from AI labs and speech specialists (for example OpenAI's transcription models, ElevenLabs Scribe, Deepgram, AssemblyAI and the cloud providers' speech services), typically with diarisation, timestamps and vocabulary features.
- Multimodal models that accept audio directly (such as Gemini), which can transcribe and also answer questions, summarise or extract from the audio in one call.
- Open-weight models you can self-host for privacy or cost reasons.
Model line-ups change quickly. For example, in August 2026 OpenAI announced that several of its transcription models (including whisper-1 and the gpt-4o transcribe family) will be removed from its API on 26 February 2027, with newer transcription models as the migration path. Always check deprecation pages and keep your pipeline provider-agnostic.
Hands-on: compute WER and entity accuracy
# pip install jiwer
import re, jiwer
def normalise(t: str) -> str:
t = t.lower()
t = re.sub(r"[^\w\s']", " ", t) # drop punctuation
return re.sub(r"\s+", " ", t).strip()
reference = "send the invoice to Amina by Friday for fifteen thousand rupees"
hypothesis = "send an invoice to Amna Friday for fifty thousand rupees"
wer = jiwer.wer(normalise(reference), normalise(hypothesis))
print(f"WER: {wer:.0%}")
ENTITIES = ["amina", "friday", "fifteen thousand"]
hyp = normalise(hypothesis)
missed = [e for e in ENTITIES if e not in hyp]
print("Entity accuracy:", f"{1 - len(missed) / len(ENTITIES):.0%}", "missed:", missed)This small example shows why entity checks matter: "fifteen" versus "fifty" is one substitution in WER terms and a costly mistake in business terms.
Prompting audio-capable models
When a multimodal model takes audio directly, give it the same briefing you would give a human transcriber:
Transcribe this 20-minute sales call between our rep (Sara) and a client
(Mr Qureshi). Label speakers by name. The call mixes English and Urdu:
transcribe Urdu in Urdu script and add an English translation in brackets.
Keep numbers as digits. Mark inaudible sections as [inaudible mm:ss].
Vocabulary: Optimize All, Shopify, CRO, ROAS.Then keep the raw transcript and treat any summary or extraction as a separate, checkable step.
Going further
Build an evaluation set of 20 to 50 audio clips reflecting your real conditions: accents, languages, noise, jargon. Create reference transcripts carefully (this is labour-intensive, but essential). Measure WER overall and per segment (speaker group, noise level), plus field-level accuracy for the entities you care about. Re-run whenever you change provider or settings.
Key takeaways
- STT turns speech into text; look for language and code-switching support, diarisation, timestamps, custom vocabulary, streaming vs batch.
- Word error rate is the standard metric, but also measure accuracy on names, numbers and key entities.
- Common failures: names, numbers, overlapping speech, uneven accuracy across accents, noise and hallucinated text.
- Fix audio first, supply vocabulary, post-process carefully, and respect consent and retention rules.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Record or collect five short clips representative of your use case, write reference transcripts, and compute WER and name/number accuracy for one STT tool.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.