Emerging Tech Horizons: What's Next After Today's AIFrontier AI trends · Lesson 4 of 16

Multimodal AI, world models and on-device AI

Article · 7 min · 8 min lecture

Video lecture

Multimodal AI, world models and on-device AI

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Beyond text

  • Multimodal AI
  • World models
  • On-device AI
  • Now vs outlook

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Beyond text

Three further trends broaden where AI can be used: multimodality (models that understand and generate text, images, audio and video together), world models (models that learn how environments behave and can simulate them), and on-device AI (models running locally on phones, laptops and edge hardware). Each opens specific opportunities and has specific limits.

Multimodal AI: one model, many senses

Modern frontier models accept images, documents, screenshots, audio and increasingly video as input, and many can produce images, speech and video. Real-time voice interfaces allow natural spoken conversation with low latency, including interruptions.

Practical uses now: reading invoices, forms and receipts; analysing charts and dashboards from screenshots; inspecting product photos for defects; transcribing and summarising calls; generating product imagery and short video variations for ads (with disclosure and rights checks); voice agents for bookings and support.

Limits: fine detail in dense images, precise counting and measurement, spatial reasoning, and hallucinated details in generated media. Synthetic media also raises provenance and trust issues (module five).

World models: learning how the world behaves

A world model learns to predict how an environment changes in response to actions, so it can simulate "what happens if". This matters for robotics, autonomous vehicles, games and training agents safely in simulation. Examples include Google DeepMind's Genie 3 (announced August 2025), which generates interactive environments from text prompts that can be navigated in real time for short periods, and NVIDIA's Cosmos platform of world foundation models aimed at physical AI development.

Outlook (reasoned): world models could make training robots and agents cheaper and safer by moving more learning into simulation, and could power new forms of interactive media and design tools. The open questions are fidelity (how closely simulations match reality), consistency over long periods, and compute cost. For most businesses, world models are a watch item today, and an experiment item for those in gaming, simulation, robotics, architecture or training.

On-device AI: private, fast, offline

Phones and laptops now ship with neural processing units (NPUs) designed to run AI models locally. Platform vendors offer on-device models and frameworks (for example, Apple's on-device foundation models with a developer framework, Google's Gemini Nano on supported Android devices, and Windows PCs with NPUs). Small language models, compressed and optimised versions of larger ones, can run offline for tasks like summarisation, transcription, classification and smart replies.

Why it matters:

  • Privacy: sensitive data can stay on the device.
  • Latency: no network round trip.
  • Cost: no per-call cloud fees.
  • Offline: field teams, rural areas, flights, secure sites.

Trade-offs: smaller models are less capable; devices vary widely in hardware; updates and model management across a fleet are harder; battery and heat constraints apply. The common pattern is hybrid: simple and private tasks on the device, harder tasks escalated to the cloud with user consent.

Hands-on: a multimodal opportunity scan

Walk through a typical week in one team and list every moment where information arrives as an image, PDF, audio or video. Then use this prompt with any multimodal assistant to prototype one.

You are helping me prototype a multimodal workflow. I will upload {5 sample receipts / 3 call recordings / 10 product photos}.
1. Extract {fields: supplier, date, total, VAT} into a table.
2. Flag any item where you are uncertain, and say why.
3. List what would need to be true for this to work reliably at 500 items per week
   (image quality, formats, languages such as Arabic or Urdu, edge cases).
Do not guess unreadable values; mark them UNREADABLE.

Check every extracted value against the originals to measure accuracy before drawing conclusions.

Worked example: a Lahore field-sales team

A consumer-goods distributor in Lahore wanted reps to log shelf photos and orders in shops with poor connectivity. A pilot used on-device transcription and summarisation on reps' phones for voice notes in Urdu and English, and queued shelf photos for cloud analysis when a connection returned. Accuracy on Urdu voice notes varied by device and noise level, so the team added a quick confirm-or-edit step. The privacy benefit (voice notes stayed on the device until summarised) helped with shop owners' trust.

Pitfalls

  • Assuming a multimodal demo on clean images will work on real, messy photos.
  • Treating world models as production tools for most businesses today.
  • Ignoring device diversity when planning on-device features.

How to measure success

Accuracy on real samples in your languages and conditions, processing time, cost per item, and adoption by the people who use it. For on-device features, measure performance across the lowest-spec devices you support.

Key takeaways

  • Multimodal models read and generate images, audio and video, enabling document, visual inspection and voice use cases, with limits on fine detail and precision.
  • World models simulate how environments respond to actions; they matter most for robotics, simulation and games and are a watch item for most businesses.
  • On-device AI offers privacy, low latency, lower cost and offline use; hybrid device-plus-cloud is the common pattern.
  • Test on real, messy samples in your languages and on your lowest-spec devices before scaling.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which use case best suits on-device AI?
  2. What is a world model?
  3. A receipt-extraction demo worked perfectly on clean scans. What should you do before rollout?

Put it into practice

Run the multimodal opportunity scan for one team and prototype one workflow with the prompt, checking every extracted value against originals.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.