Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsEdge AI, privacy and honest benchmarking · Lesson 12 of 16
On-device AI and small language models
Video lecture
On-device AI and small language models
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 On-device AI and small models
Here is a quiet revolution. The phone in your pocket can now run a language model that rewrites messages, classifies requests and extracts data with no network, no per-request cost and no data leaving the device. In this lesson you will meet the small model families, the runtimes on Apple, Android, browsers and Windows, the design patterns that make small models shine, and how to evaluate them on real devices.
0:30 Analogy: toolbox
Think of small models as specialist tools in a toolbox. A big frontier model is a skilled general contractor who can do almost anything, but you pay for every hour and they have to travel to your site. A small on-device model is a good screwdriver: it lives in your pocket, costs nothing each time you use it, and does one job very well. You would not ask a screwdriver to build a house. But you also would not call a contractor to tighten one screw.
1:07 Small models
Small language models, roughly half a billion to eight billion parameters, trade breadth of knowledge and deep reasoning for four things: privacy, zero marginal cost, offline operation and very low latency. For narrow jobs, classification, extraction, short rewrites, autocomplete, routing, a well-chosen small model is often all you need. Current families include Gemma 4's edge sizes, Mistral's Ministral 3, Microsoft's Phi-4-mini, small Qwen models, and at the top of edge territory, gpt-oss twenty B. As always, check the model cards.
1:42 Runtimes
Where do they run? On Apple devices, MLX runs open models on Apple silicon, and Apple's Foundation Models framework gives apps access to the on-device model behind Apple Intelligence. On Android, Google offers Gemini Nano through ML Kit on supported devices, and LiteRT and MediaPipe run open models like Gemma. In the browser, Chrome has built-in AI APIs backed by Gemini Nano, and WebLLM and Transformers dot js run open models on WebGPU. On Windows, there is ONNX Runtime and Microsoft's Foundry Local. Availability shifts, so check current docs.
2:21 Built-in vs your own
Built-in models versus shipping your own. A built-in model means your users download nothing extra. But you do not control the version, and it depends on device and region. Shipping an open model gives you control and consistency, at the cost of app size and update management. Many apps use both: the built-in model where available, their own small model elsewhere, and a cloud fallback with user consent.
2:51 Five patterns
Five patterns make small models shine. One, on-device first, with a cloud fallback when confidence is low, and the user's permission. Two, narrow the task: tight instructions, a fixed label set, constrained output. Three, retrieval, not memorization: pair the model with a local search index. Four, fine-tune for the job, because a small model trained on a few thousand examples of one task can beat a much bigger general model on that task. Five, respect battery and heat.
3:25 Worked example: field sales in KSA
A worked example. A beverage distributor in Saudi Arabia sends reps to shops in remote areas with patchy coverage. Their tablet app turns voice-note transcripts into structured visit reports, shop, stock issues, next order, in Arabic and English, using a small model fine-tuned on three thousand reviewed reports. It runs offline in seconds and syncs when the tablet reconnects. Big analytical questions go to a cloud assistant back at head office.
3:56 Hands-on + evaluation
The hands-on uses MLX on a Mac: one command to generate text with a four-bit small model, and a tiny Python classifier that forces one label from a fixed list and maps anything unexpected to other. Never trust free text from a small model. Then evaluate properly. Test on the devices your users actually own, often mid-range phones, over a twenty-minute session. Measure latency, memory, battery drain and heat, plus task accuracy on a gold set.
4:29 Simple example: note tagging
A simple example. A note-taking app wants to suggest three tags for every note a user writes. Tags come from the user's own existing list. That is a narrow task, a fixed label set, short input, and mistakes are harmless because the user can remove a tag with one tap. A small model on the phone does it instantly, offline, for free, and the note never leaves the device. This is the sweet spot for on-device AI.
5:02 Shipping models in apps
Shipping a model inside an app raises practical questions. Model files can be several gigabytes, so download them after install, on Wi-Fi, with a progress indicator and a resume option. Check free storage first. Plan updates: a new model version is a new download, so version it and allow rollback. And remember older devices. In many markets, including large parts of Pakistan and the wider region, mid-range and older phones dominate, so decide a minimum device spec and give everyone else the cloud fallback with consent.
5:39 FAQ
A question from product teams: how do we update an on-device model safely? Treat it like any app release. Version the model file, ship it to a small percentage of users first, compare accuracy and crash or battery metrics against the previous version, and keep the ability to roll back. Also keep the app working if the new model has not downloaded yet. And a second question: can small models handle Arabic or Urdu? Some do reasonably, many do not. Test your languages early, because language coverage is often the deciding factor for small models.
6:20 Try this now
Try this now. Look at your product or your daily work and list three tasks that are narrow, frequent and low-risk: things like tagging, short rewrites or routing. For each one, write the fixed set of outputs it should produce. If you can write that set in one line, the task is probably a good fit for a small model.
6:46 Watch me do it
Watch me do it. On a Mac, I install mlx lm and run a four-bit, four billion parameter model with a one-line prompt to rewrite a message politely. It answers in about a second. Now a real task: classify thirty support messages into billing, delivery, technical or other. I run the Python classifier from the lesson, which forces one label and maps anything unexpected to other. It gets twenty-six right. Three misses are messages mentioning two problems; I add a priority rule to the system prompt and it gets twenty-eight. Then I check the device side: I watch memory use in the activity monitor, around three gigabytes, and run the classifier on a loop for twenty minutes while noting battery and temperature. Everything stays reasonable. Last step, I note what would happen on an older device, and decide the cloud fallback threshold for users below our minimum spec.
7:50 Recap
Recap. Small models win on privacy, cost, offline use and latency for narrow tasks. Choose between built-in platform models and your own, and often use both. Narrow the task, add local retrieval, fine-tune when needed, and always test on real target devices. Your next step: pick one narrow task from your product, run a small model on your own device, and measure accuracy on thirty examples plus latency and memory.
Why small models matter
Small language models (SLMs), roughly 0.5–8B parameters, run on phones, laptops, browsers, cars and factory gateways. They trade breadth of knowledge and deep reasoning for privacy, zero marginal cost, offline operation and very low latency. For many narrow tasks (classification, extraction, short rewriting, autocomplete, on-device search, routing) a well-chosen small model, sometimes fine-tuned, is all you need.
Current small-model families (September 2026 snapshot; verify on model cards)
- Gemma 4 edge-sized variants (Apache-2.0), designed for phones and laptops.
- Ministral 3 (3B, 8B, 14B) from Mistral, Apache-2.0, with base, instruct and reasoning variants and image understanding.
- Phi-4-mini and related Phi-4 models (MIT), strong reasoning for their size.
- Qwen3 / Qwen3.5 small dense sizes (mostly Apache-2.0), good multilingual coverage.
- gpt-oss-20b sits at the top of "edge" territory: an MoE with ~3.6B active parameters that runs within 16 GB of memory.
Platform runtimes and built-in models
| Platform | Options (check current docs) |
|---|---|
| Apple (macOS, iOS) | MLX / mlx-lm for running open models on Apple silicon; Apple's Foundation Models framework gives apps access to the on-device model that powers Apple Intelligence on supported devices |
| Android | Google's Gemini Nano via ML Kit GenAI APIs on supported devices; LiteRT / MediaPipe LLM Inference for running open models such as Gemma on-device |
| Browser | Chrome's built-in AI APIs (for example Prompt, Summarizer and Translator APIs backed by Gemini Nano, availability varies by API and platform); WebLLM and Transformers.js run open models via WebGPU/WebAssembly |
| Windows / cross-platform | ONNX Runtime GenAI, Foundry Local from Microsoft, llama.cpp, Ollama, LM Studio |
| Linux edge gateways | llama.cpp, Ollama, ONNX Runtime; NVIDIA Jetson-class devices for GPU at the edge |
Built-in OS models are attractive because the user downloads nothing extra, but you do not control the model version, and availability depends on device and region. Open models you ship give control at the cost of app size and update management.
Design patterns for on-device AI
- On-device first, cloud fallback. Try the local model; if confidence is low or the task is complex, ask the user's permission to use a cloud model.
- Narrow the task. Small models shine with tight instructions, constrained output and a fixed label set. Do not ask a 3B model to be a general assistant.
- Retrieval, not memorization. Pair a small model with a local search index over the user's own data.
- Fine-tune for the job. A small model fine-tuned on a few thousand examples of one task often beats a much larger general model on that task (next course).
- Respect the battery and the heat. Batch work while charging; keep prompts short; unload models when idle.
Hands-on: run a small model with MLX on a Mac
pip install mlx-lm
# Model IDs change; browse the mlx-community organization on Hugging Face for current conversions
mlx_lm.generate --model mlx-community/Qwen3-4B-4bit \
--prompt "Rewrite politely in one sentence: 'send the invoice now'" --max-tokens 60And in Python, a narrow classifier with a fixed label set:
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/Qwen3-4B-4bit")
LABELS = ["billing", "delivery", "technical", "other"]
def classify(text: str) -> str:
messages = [{"role": "system", "content": f"Reply with exactly one label from {LABELS}. No other words."},
{"role": "user", "content": text}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
out = generate(model, tokenizer, prompt=prompt, max_tokens=5).strip().lower()
return out if out in LABELS else "other" # never trust free text; map unknowns
print(classify("My parcel has been stuck in Dubai customs for 5 days"))Some reasoning models emit a thinking block before the answer; check the model card for how to disable or strip it for classification.
Worked example: a field-sales app in Saudi Arabia
A beverage distributor's reps visit shops across remote regions with patchy coverage. Their tablet app uses an on-device small model to turn voice-note transcripts into structured visit reports (shop, stock issues, next order) in Arabic and English. Reports sync when the tablet reconnects. The team fine-tuned a small model on 3,000 reviewed reports; it runs offline in seconds. Complex questions ("why are sales down in this region?") go to a cloud analytics assistant back at the office.
Evaluating small models
- Test on the exact device class your users have; mid-range phones differ hugely from flagships.
- Measure latency, memory, battery drain and thermal throttling over a 20-minute session, not just a single request.
- Evaluate quality on your narrow task with a gold set; small models can be excellent on narrow tasks and poor outside them.
Pitfalls
- Using an SLM as a general chatbot and being disappointed.
- Shipping a 4 GB model inside an app without a download strategy.
- Ignoring older devices in markets where they dominate.
- Assuming "on-device" means compliant: logs and sync may still move data.
How to measure success
On your target devices, the small model meets your task accuracy bar, p95 latency and battery budget, and falls back gracefully when it cannot.
Key takeaways
- Small models (≈0.5–8B) win on privacy, cost, offline use and latency for narrow tasks
- Options include open SLMs (Gemma 4 edge, Ministral 3, Phi-4-mini, small Qwen) and OS built-in models
- Patterns: on-device first with cloud fallback, narrow tasks, local retrieval, task fine-tuning
- Evaluate on real target devices: latency, memory, battery, heat and task accuracy
- On-device is not automatically compliant: check logs and sync
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Pick one narrow task from your product, run a small model for it on your own device, and measure accuracy on 30 examples plus latency and memory.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.