Voice AI & Conversational AgentsDesigning voice conversations · Lesson 10 of 17
Multilingual voice: Urdu, Arabic and English
Video lecture
Multilingual voice: Urdu, Arabic and English
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Multilingual voice
Listen to how people actually talk in Karachi: mera order abhi tak nahi aaya, can you check the status? Or in Dubai, where a caller slides between Gulf Arabic and English in one sentence. In Pakistan, the UAE and Saudi Arabia, multilingual and code-switched speech is the default, not the exception. An agent designed for one language at a time will fail a lot of your callers. In this lesson, you'll learn where multilingual voice breaks, how to handle scripts, numbers and dates, how to pick register and politeness, and how to evaluate per language.
0:41 Why multilingual by design
Why is multilingual design a must in these markets? Because your callers don't speak one language at a time. An agent that only works well in English quietly fails a large share of real conversations in Pakistan and the Gulf, and those failures often go unnoticed in overall averages. Designing for code-switching, dialects and scripts from day one is far cheaper than retrofitting it after complaints.
1:10 Where it breaks
Problems appear at every layer. Speech recognition struggles with dialects, like Gulf, Najdi, Egyptian or Levantine Arabic, regional Urdu accents, code-switching, names and numbers. The fix: test on real recordings from your callers, pick models strong in your dialects, and allow keypad input for numbers. Language detection gets confused by short words like okay or haan, so detect on longer turns and avoid flip-flopping. The language model needs clear instructions on which language and register to reply in. And text to speech needs voices tested in each language, with pronunciation help for names and brands.
1:51 Analogy: a multilingual front desk
An analogy: multilingual voice design is like staffing a front desk in a busy mall in Dubai or Karachi. The best receptionists switch languages naturally, use the English word when it's clearer, greet people appropriately, and don't make a customer choose a language from a menu before they can speak. But the official notices on the wall are professionally translated and checked. That's the balance: natural, flexible conversation, and fixed, reviewed wording for anything legal.
2:24 Scripts and formats
Watch scripts and formats. Urdu is written in Perso-Arabic script, right to left, but Roman Urdu, in Latin letters, is common in messages, and recognition may output either. Decide what your systems expect. Arabic speakers may use Western or Eastern Arabic digits, so normalize before calling tools. Callers may mention dates by the Hijri calendar, especially around Ramadan and Eid, so resolve carefully and confirm the Gregorian date. And names are transliterated many ways, like Mohammed, Muhammad or Mohamed, so match flexibly in CRM lookups.
3:01 Register and politeness
Register and politeness matter as much as accuracy. For spoken service in the UAE and Saudi Arabia, many businesses use a friendly, lightly formal Gulf style. Modern Standard Arabic can sound stiff on the phone, but it's safer for legal statements. Decide deliberately and test with local listeners. In Urdu, the respectful aap form is standard, and English technical terms like OTP, account or order are often clearer than pure Urdu. And greetings count: assalam o alaikum, with the right response.
3:36 Code-switching strategies
You have three code-switching strategies. Mirror: reply in the caller's dominant language, allowing common English terms. That's the most natural. Ask once: would you prefer Urdu or English? Simple, but slightly robotic. Or primary with switching: default to Arabic in Saudi Arabia and switch to English when detected. Some platforms offer special handling for mixed-language speech and language detection tools. Test what works on your callers. And one rule holds regardless: AI disclosures, recording notices and legal statements use approved, native-reviewed translations. Never let the model improvise them.
4:14 Worked example: Pakistani telecom
A Pakistani telecom built a bill-inquiry agent. It greets in Urdu with an English option, and introduces itself as an AI assistant. The team tested speech recognition on two hundred real call snippets from Karachi, Lahore and Peshawar, chose the model with the fewest errors on numbers and names, and added keypad entry for account numbers. The prompt allows English words like bill, package and data inside Urdu replies. One voice handles both languages, with a pronunciation dictionary for package names. And success is measured separately for each language, to catch hidden gaps.
4:54 Evaluate per language
The lesson text includes a starter multilingual evaluation set, with Urdu from Karachi, Lahore and Peshawar, Gulf and Najdi Arabic, English, and code-switched lines, each with the expected intent and reply language. Record them, or use real consented samples, run them through your agent, and score intent accuracy, reply language and task success per language and dialect. Never report a single overall success rate. It can hide a language that's failing badly.
5:25 Example 2: Dubai clinic dialects
A second example. A Dubai clinic's agent handles Arabic and English. Testing showed Egyptian and Levantine Arabic speakers were misunderstood more often than Gulf speakers, because the chosen speech model was tuned toward Gulf speech in its test set. The team added test recordings from each dialect, compared two recognition models, and switched to the one with more balanced accuracy. They also added a gentle fallback: if recognition confidence is low twice, the agent offers English or a human. Success rates evened out across dialects.
6:02 Common mistakes
Common multilingual mistakes. Testing only in English and assuming other languages just work. Letting the model translate mandatory legal statements on the fly. Flip-flopping languages on short words like okay or haan. And reporting a single overall success rate that hides a language or dialect that's failing badly. Measure per language, and fix the weakest first.
6:26 Watch me do it: per-language testing
Watch me do it. I take the multilingual evaluation set and adapt it for Nova Dental: fifteen utterances, five each in Gulf Arabic, Egyptian Arabic and English, plus three code-switched lines like hi, abi ahjiz appointment for cleaning. I ask colleagues from different backgrounds to record them on their phones, with consent to use them for testing. I run all of them through the agent and fill a results table: intent correct, reply language correct, task completed. English is fine. Gulf Arabic is mostly fine. Egyptian Arabic fails more often, mostly on dates. So I make three changes in the config. In the prompt, I add: reply in the language the caller uses; for Arabic, use friendly, clear Modern Standard Arabic with common Gulf expressions. I enable the language detection system tool, restricted to switching early in the call to avoid flip-flopping. And I replace the disclosure and recording lines in Arabic with the version our native-speaking colleague approved, fixed in the first message rather than generated. Then I rerun the full set and report success per language, never as one average. Egyptian Arabic improves but remains the weakest, so I note it as the top item for next week, with a plan to compare a second speech recognition model.
7:57 Recap and next step
Recap. In our markets, multilingual and code-switched speech is normal. Test recognition on real dialects, handle scripts, digits and Hijri dates, choose register deliberately, pick a code-switching strategy, and use approved translations for mandatory statements. Your next step: adapt the evaluation set to your callers, record at least five utterances per language or dialect, run them, and compare success rates. Fix the weakest language first.
Multilingual is the default in our markets
In Pakistan, the UAE and Saudi Arabia, many callers switch languages mid-sentence. A Karachi customer might say "Mera order abhi tak nahi aaya, can you check the status?" A Dubai caller might mix Gulf Arabic and English. A Riyadh caller might speak Najdi Arabic, while your content is in Modern Standard Arabic. Designing for one language at a time fails these callers.
Challenges by layer
| Layer | Challenge | Mitigation |
|---|---|---|
| ASR | Dialects (Gulf, Najdi, Egyptian, Levantine; Urdu regional accents), code-switching, names, numbers | Test ASR on real recordings from your callers; choose models strong in your dialects; add keyword/vocabulary hints where supported; allow keypad input for numbers |
| Language detection | Short utterances ("ok", "yes", "haan") are ambiguous | Detect on longer turns; confirm language switch; avoid flip-flopping (some platforms can restrict switching to the start of the call) |
| LLM | Understanding mixed input; responding in the right register | Instruct "reply in the language the caller uses"; provide examples in each language; specify MSA vs dialect style |
| TTS | Voice quality in each language; pronunciation of names, brands, English words inside Urdu or Arabic sentences | Pick voices tested in each language; use multilingual models; pronunciation dictionaries; consider separate voices per language |
| Content | Knowledge base only in English | Maintain approved translations; retrieve in the caller's language or translate carefully with review |
Script and formatting issues
- Urdu is written in Nastaliq-style Perso-Arabic script, right to left; Roman Urdu (Urdu in Latin letters) is common in messages. ASR may output either; decide what your downstream systems expect.
- Arabic numerals: Western digits (0-9) vs Eastern Arabic digits. Normalize before tools.
- Dates: callers may reference the Hijri calendar, especially around Ramadan and Eid. Resolve carefully and confirm Gregorian dates.
- Names: transliteration varies (Mohammed, Muhammad, Mohamed). Match flexibly in CRM lookups.
Register and politeness
- Arabic: many businesses use a friendly, lightly formal Gulf style for spoken service in the UAE and KSA; MSA can sound stiff on the phone but is safer for legal statements. Decide deliberately and test with local listeners.
- Urdu: respectful forms ("aap") are standard for customers; English technical terms (OTP, account, order) are often clearer than pure Urdu equivalents.
- Greetings matter: "Assalam o alaikum" / "As-salamu alaykum" and appropriate responses.
Code-switching strategy
Options:
- Mirror: reply in the caller's dominant language, allowing common English terms. Most natural.
- Ask once: "Would you prefer Urdu or English?" at the start. Simple, slightly robotic.
- Primary language with switching: default to Arabic in KSA, switch to English on detection.
Some platforms offer specific modes for mixed-language speech (for example handling of Hindi-English mixing) and language detection system tools. Test what works on your callers' speech.
Legal and disclosure in multiple languages
AI disclosure, recording notices and key terms should be delivered in the caller's language. Keep approved translations of mandatory statements, reviewed by native speakers, and do not let the LLM improvise them.
Worked example: a Pakistani telecom's bill-inquiry agent
- Default greeting in Urdu with an English option: "Assalam o alaikum, main [brand] ki AI assistant hoon. Aap Urdu ya English mein baat kar sakte hain."
- ASR tested on 200 real call snippets from Karachi, Lahore and Peshawar; the team chose the model with the lowest error on numbers and names and enabled keypad entry for account numbers.
- Prompt allows English terms ("bill", "package", "data") inside Urdu replies.
- TTS: one voice strong in both languages; pronunciation dictionary for package names.
- Evaluation: separate success metrics per language to detect hidden gaps.
Hands-on: a multilingual evaluation set
id,language,dialect,utterance,expected_intent,expected_language_reply,notes
u01,ur,karachi,"Mera bill zyada aaya hai, check kar dein",bill_inquiry,ur,
u02,ur-en,lahore,"Mujhe apna data package change karna hai please",package_change,ur,code-switch
u03,ar,gulf,"أبغى أغير موعدي لبكرة العصر",reschedule,ar,tomorrow afternoon
u04,ar,najdi,"وش أوقات الدوام يوم الجمعة؟",hours_query,ar,
u05,en,uae,"Hi, can I move my appointment to Thursday?",reschedule,en,
u06,ar-en,gulf,"Hi, أبي أحجز appointment للأسنان",book,ar,code-switch
u07,ur,peshawar,"Account number hai 0 3 4 5... 6 7 8... 9 0 1 2",identify,ur,digits with pausesRecord these (or real, consented samples) as audio, run them through your agent, and score intent accuracy, reply language and task success per language and dialect.
Pitfalls
- Testing only in English and assuming other languages "just work".
- Letting the LLM translate mandatory legal statements on the fly.
- Reporting a single overall success rate that hides a failing language.
Key takeaways
- Code-switched Urdu-English and Arabic-English speech is normal in PK, UAE and KSA; design and test for it explicitly.
- Test ASR on real dialect recordings, allow keypad input for numbers, and avoid language flip-flopping on short utterances.
- Normalize scripts and digits, confirm Hijri-referenced dates, and match transliterated names flexibly.
- Use approved, native-reviewed translations for disclosures and legal statements; report success per language and dialect.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Adapt the multilingual evaluation set to your callers. Record at least five utterances per language or dialect, run them through your agent, and report success per language. Fix the weakest language first.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.