Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsEdge AI, privacy and honest benchmarking · Lesson 12 of 16

On-device AI and small language models

Article · 15 min · 8 min lecture

Video lecture

On-device AI and small language models

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

On-device AI and small models

  • Families and runtimes
  • Design patterns
  • Evaluating on real devices

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why small models matter

Small language models (SLMs), roughly 0.5–8B parameters, run on phones, laptops, browsers, cars and factory gateways. They trade breadth of knowledge and deep reasoning for privacy, zero marginal cost, offline operation and very low latency. For many narrow tasks (classification, extraction, short rewriting, autocomplete, on-device search, routing) a well-chosen small model, sometimes fine-tuned, is all you need.

Current small-model families (September 2026 snapshot; verify on model cards)

  • Gemma 4 edge-sized variants (Apache-2.0), designed for phones and laptops.
  • Ministral 3 (3B, 8B, 14B) from Mistral, Apache-2.0, with base, instruct and reasoning variants and image understanding.
  • Phi-4-mini and related Phi-4 models (MIT), strong reasoning for their size.
  • Qwen3 / Qwen3.5 small dense sizes (mostly Apache-2.0), good multilingual coverage.
  • gpt-oss-20b sits at the top of "edge" territory: an MoE with ~3.6B active parameters that runs within 16 GB of memory.

Platform runtimes and built-in models

PlatformOptions (check current docs)
Apple (macOS, iOS)MLX / mlx-lm for running open models on Apple silicon; Apple's Foundation Models framework gives apps access to the on-device model that powers Apple Intelligence on supported devices
AndroidGoogle's Gemini Nano via ML Kit GenAI APIs on supported devices; LiteRT / MediaPipe LLM Inference for running open models such as Gemma on-device
BrowserChrome's built-in AI APIs (for example Prompt, Summarizer and Translator APIs backed by Gemini Nano, availability varies by API and platform); WebLLM and Transformers.js run open models via WebGPU/WebAssembly
Windows / cross-platformONNX Runtime GenAI, Foundry Local from Microsoft, llama.cpp, Ollama, LM Studio
Linux edge gatewaysllama.cpp, Ollama, ONNX Runtime; NVIDIA Jetson-class devices for GPU at the edge

Built-in OS models are attractive because the user downloads nothing extra, but you do not control the model version, and availability depends on device and region. Open models you ship give control at the cost of app size and update management.

Design patterns for on-device AI

  1. On-device first, cloud fallback. Try the local model; if confidence is low or the task is complex, ask the user's permission to use a cloud model.
  2. Narrow the task. Small models shine with tight instructions, constrained output and a fixed label set. Do not ask a 3B model to be a general assistant.
  3. Retrieval, not memorization. Pair a small model with a local search index over the user's own data.
  4. Fine-tune for the job. A small model fine-tuned on a few thousand examples of one task often beats a much larger general model on that task (next course).
  5. Respect the battery and the heat. Batch work while charging; keep prompts short; unload models when idle.

Hands-on: run a small model with MLX on a Mac

pip install mlx-lm
# Model IDs change; browse the mlx-community organization on Hugging Face for current conversions
mlx_lm.generate --model mlx-community/Qwen3-4B-4bit \
  --prompt "Rewrite politely in one sentence: 'send the invoice now'" --max-tokens 60

And in Python, a narrow classifier with a fixed label set:

from mlx_lm import load, generate

model, tokenizer = load("mlx-community/Qwen3-4B-4bit")
LABELS = ["billing", "delivery", "technical", "other"]

def classify(text: str) -> str:
    messages = [{"role": "system", "content": f"Reply with exactly one label from {LABELS}. No other words."},
                {"role": "user", "content": text}]
    prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
    out = generate(model, tokenizer, prompt=prompt, max_tokens=5).strip().lower()
    return out if out in LABELS else "other"   # never trust free text; map unknowns

print(classify("My parcel has been stuck in Dubai customs for 5 days"))

Some reasoning models emit a thinking block before the answer; check the model card for how to disable or strip it for classification.

Worked example: a field-sales app in Saudi Arabia

A beverage distributor's reps visit shops across remote regions with patchy coverage. Their tablet app uses an on-device small model to turn voice-note transcripts into structured visit reports (shop, stock issues, next order) in Arabic and English. Reports sync when the tablet reconnects. The team fine-tuned a small model on 3,000 reviewed reports; it runs offline in seconds. Complex questions ("why are sales down in this region?") go to a cloud analytics assistant back at the office.

Evaluating small models

  • Test on the exact device class your users have; mid-range phones differ hugely from flagships.
  • Measure latency, memory, battery drain and thermal throttling over a 20-minute session, not just a single request.
  • Evaluate quality on your narrow task with a gold set; small models can be excellent on narrow tasks and poor outside them.

Pitfalls

  • Using an SLM as a general chatbot and being disappointed.
  • Shipping a 4 GB model inside an app without a download strategy.
  • Ignoring older devices in markets where they dominate.
  • Assuming "on-device" means compliant: logs and sync may still move data.

How to measure success

On your target devices, the small model meets your task accuracy bar, p95 latency and battery budget, and falls back gracefully when it cannot.

Key takeaways

  • Small models (≈0.5–8B) win on privacy, cost, offline use and latency for narrow tasks
  • Options include open SLMs (Gemma 4 edge, Ministral 3, Phi-4-mini, small Qwen) and OS built-in models
  • Patterns: on-device first with cloud fallback, narrow tasks, local retrieval, task fine-tuning
  • Evaluate on real target devices: latency, memory, battery, heat and task accuracy
  • On-device is not automatically compliant: check logs and sync

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which task best suits a 3–4B on-device model?
  2. What is a trade-off of relying on an OS built-in model (e.g. on-device model provided by the platform)?
  3. How should you test an on-device model for a market where mid-range phones dominate?

Put it into practice

Pick one narrow task from your product, run a small model for it on your own device, and measure accuracy on 30 examples plus latency and memory.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.