Emerging Tech Horizons: What's Next After Today's AIFrontier AI trends · Lesson 4 of 16
Multimodal AI, world models and on-device AI
Video lecture
Multimodal AI, world models and on-device AI
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Beyond text
Text was just the beginning. Today's AI can look at a photo, listen to a call, watch a video, and even simulate how a world behaves. And more of it is running directly on the phone in your pocket. In this lesson you'll explore multimodal AI, world models and on-device AI, with a clear view of what works now and what's still an outlook.
0:28 Multimodal now
Multimodal first. Frontier models now accept images, documents, screenshots, audio and increasingly video, and many can generate images, speech and video too. Real-time voice lets you have a natural spoken conversation, interruptions and all. What's practical right now? Reading invoices and receipts. Analysing dashboards from screenshots. Spotting defects in product photos. Summarising calls. Generating ad image variations, with disclosure and rights checks. And voice agents for bookings and support.
0:58 Why it matters
Why does this matter? Because a huge amount of business information never arrives as neat text. It arrives as photos of receipts, screenshots of dashboards, voice notes from the field, scanned forms and video. Until recently, turning that into usable data meant manual typing. Multimodal and on-device AI change the economics of that work, especially for small teams and field operations. And world models hint at a future where more testing and training happens in simulation first. Knowing what's ready now, and what isn't, helps you pick the right experiments.
1:37 Eyes, imagination and a desk nearby
Here's an analogy. Think of multimodal AI as giving your assistant eyes and ears, not just a keyboard. World models are like giving them imagination: the ability to picture what would happen if they pushed a box or turned a corner. And on-device AI is like the assistant working at your desk instead of in a distant office: faster, more private, but with a smaller toolbox. Each adds something different, and each has its own limits.
2:10 Simple example: invoice photos
A simple example of multimodal in practice. A small restaurant in Dubai photographs supplier invoices on a phone at the back door. A multimodal assistant extracts supplier, date, items and total into a spreadsheet, and flags two invoices where the total was hard to read. The owner checks those two by hand. What used to be an evening of typing becomes ten minutes of checking. But notice the flags: the value comes from the model being honest about uncertainty.
2:44 Limits
But know the limits. Multimodal models still struggle with fine detail in dense images, precise counting and measurement, and some spatial reasoning. Generated media can contain invented details. And a demo on clean, well-lit images tells you little about messy real-world photos taken on a busy shop floor. Always test on your own samples, in your own languages, like Arabic and Urdu.
3:11 World models
Now world models. A world model learns how an environment changes when you act in it, so it can simulate what happens if. Google DeepMind's Genie 3, announced in August 2025, generates interactive environments from a text prompt that you can move through in real time for short periods. NVIDIA's Cosmos platform offers world foundation models for building physical AI, like robots and autonomous vehicles.
3:39 World models outlook
Why could this matter? As outlook, not prediction: world models could make training robots and agents cheaper and safer by moving more learning into simulation, and could power new interactive media and design tools. The open questions are fidelity to reality, consistency over longer periods, and compute cost. For most businesses, world models are a watch item today. For gaming, simulation, robotics, architecture and training, they're worth an experiment.
4:09 On-device AI
Now on-device AI. Phones and laptops ship with neural processing units built to run models locally. Apple offers on-device foundation models to developers, Google runs Gemini Nano on supported Android phones, and many Windows PCs include NPUs. Small language models can summarise, transcribe, classify and suggest replies without a connection. The benefits: privacy, because data can stay on the device; speed, with no network round trip; lower cost; and offline use for field teams and secure sites.
4:42 Trade-offs and the hybrid pattern
The trade-offs are real. Small models are less capable. Devices vary widely, so a feature that flies on a flagship phone may crawl on an older one. Managing models across a fleet is harder, and battery and heat matter. That's why the common pattern is hybrid: simple and private tasks on the device, harder ones escalated to the cloud with the user's consent.
5:09 Worked example: field sales
A distributor in Lahore put this together for its field-sales reps. Voice notes in Urdu and English were transcribed and summarised on the phone, even in shops with no signal, and shelf photos queued for cloud analysis when a connection returned. Accuracy varied by device and background noise, so the team added a quick confirm-or-edit step. Keeping voice notes on the device until summarised also helped build trust with shop owners.
5:40 Voice agents
A quick word on voice. Real-time voice agents can now hold natural conversations, handle interruptions and switch languages. That's opening up bookings, order status and first-line support by phone. But voice raises specific duties: tell callers they're speaking with AI where rules or good practice require it, get consent before recording, and give a quick route to a human. And test accents and dialects your customers actually use.
6:10 Three mistakes
Three common mistakes. First, assuming a demo on clean, well-lit images will work on real photos taken in a hurry. Test on your own messy samples. Second, treating world models as production tools for ordinary businesses today, rather than a watch item. Third, planning on-device features around the newest flagship phone, when many of your users have older devices with far less capability.
6:37 Try this now
Try this now. Walk through one team's week and list every moment information arrives as an image, a PDF, audio or video. Pick the most frequent one. Collect five real samples, messy ones included, and run them through a multimodal assistant with the prompt in the lesson text, which asks it to extract fields and mark anything unreadable instead of guessing. Check every value against the original. Count how many were right, how many were flagged, and how many were wrong but not flagged. That last number tells you whether it's safe to scale.
7:18 Recap
To recap. Multimodal AI opens document, visual and voice use cases now, but test on real samples. World models are an outlook for most, an experiment for some. On-device AI brings privacy, speed and offline use, usually in a hybrid with the cloud. Your next step: run the multimodal opportunity scan in the lesson text for one team and prototype one workflow, checking every value. Next module: robots, spatial computing and digital twins.
Beyond text
Three further trends broaden where AI can be used: multimodality (models that understand and generate text, images, audio and video together), world models (models that learn how environments behave and can simulate them), and on-device AI (models running locally on phones, laptops and edge hardware). Each opens specific opportunities and has specific limits.
Multimodal AI: one model, many senses
Modern frontier models accept images, documents, screenshots, audio and increasingly video as input, and many can produce images, speech and video. Real-time voice interfaces allow natural spoken conversation with low latency, including interruptions.
Practical uses now: reading invoices, forms and receipts; analysing charts and dashboards from screenshots; inspecting product photos for defects; transcribing and summarising calls; generating product imagery and short video variations for ads (with disclosure and rights checks); voice agents for bookings and support.
Limits: fine detail in dense images, precise counting and measurement, spatial reasoning, and hallucinated details in generated media. Synthetic media also raises provenance and trust issues (module five).
World models: learning how the world behaves
A world model learns to predict how an environment changes in response to actions, so it can simulate "what happens if". This matters for robotics, autonomous vehicles, games and training agents safely in simulation. Examples include Google DeepMind's Genie 3 (announced August 2025), which generates interactive environments from text prompts that can be navigated in real time for short periods, and NVIDIA's Cosmos platform of world foundation models aimed at physical AI development.
Outlook (reasoned): world models could make training robots and agents cheaper and safer by moving more learning into simulation, and could power new forms of interactive media and design tools. The open questions are fidelity (how closely simulations match reality), consistency over long periods, and compute cost. For most businesses, world models are a watch item today, and an experiment item for those in gaming, simulation, robotics, architecture or training.
On-device AI: private, fast, offline
Phones and laptops now ship with neural processing units (NPUs) designed to run AI models locally. Platform vendors offer on-device models and frameworks (for example, Apple's on-device foundation models with a developer framework, Google's Gemini Nano on supported Android devices, and Windows PCs with NPUs). Small language models, compressed and optimised versions of larger ones, can run offline for tasks like summarisation, transcription, classification and smart replies.
Why it matters:
- Privacy: sensitive data can stay on the device.
- Latency: no network round trip.
- Cost: no per-call cloud fees.
- Offline: field teams, rural areas, flights, secure sites.
Trade-offs: smaller models are less capable; devices vary widely in hardware; updates and model management across a fleet are harder; battery and heat constraints apply. The common pattern is hybrid: simple and private tasks on the device, harder tasks escalated to the cloud with user consent.
Hands-on: a multimodal opportunity scan
Walk through a typical week in one team and list every moment where information arrives as an image, PDF, audio or video. Then use this prompt with any multimodal assistant to prototype one.
You are helping me prototype a multimodal workflow. I will upload {5 sample receipts / 3 call recordings / 10 product photos}.
1. Extract {fields: supplier, date, total, VAT} into a table.
2. Flag any item where you are uncertain, and say why.
3. List what would need to be true for this to work reliably at 500 items per week
(image quality, formats, languages such as Arabic or Urdu, edge cases).
Do not guess unreadable values; mark them UNREADABLE.Check every extracted value against the originals to measure accuracy before drawing conclusions.
Worked example: a Lahore field-sales team
A consumer-goods distributor in Lahore wanted reps to log shelf photos and orders in shops with poor connectivity. A pilot used on-device transcription and summarisation on reps' phones for voice notes in Urdu and English, and queued shelf photos for cloud analysis when a connection returned. Accuracy on Urdu voice notes varied by device and noise level, so the team added a quick confirm-or-edit step. The privacy benefit (voice notes stayed on the device until summarised) helped with shop owners' trust.
Pitfalls
- Assuming a multimodal demo on clean images will work on real, messy photos.
- Treating world models as production tools for most businesses today.
- Ignoring device diversity when planning on-device features.
How to measure success
Accuracy on real samples in your languages and conditions, processing time, cost per item, and adoption by the people who use it. For on-device features, measure performance across the lowest-spec devices you support.
Key takeaways
- Multimodal models read and generate images, audio and video, enabling document, visual inspection and voice use cases, with limits on fine detail and precision.
- World models simulate how environments respond to actions; they matter most for robotics, simulation and games and are a watch item for most businesses.
- On-device AI offers privacy, low latency, lower cost and offline use; hybrid device-plus-cloud is the common pattern.
- Test on real, messy samples in your languages and on your lowest-spec devices before scaling.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run the multimodal opportunity scan for one team and prototype one workflow with the prompt, checking every extracted value against originals.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.