Gemini, Microsoft Copilot, Perplexity & the AI Tool LandscapeOpen-weight models and local AI · Lesson 14 of 19

Running AI locally: a practical introduction

Article · 17 min · 8 min lecture

Video lecture

Running AI locally: a practical introduction

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Running AI locally

  • Private
  • Offline
  • Free per use
  • Practical setup

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why run AI locally?

Running a model on your own computer means prompts and documents stay on your device, it works offline, and there is no per-use cost after setup. It is ideal for private notes, drafting, summarising personal documents, tagging, coding help on private code, and learning how models work. The trade-off: local models are usually smaller than frontier hosted models, so they are weaker at complex reasoning and have no built-in web search.

What you need

  • Hardware: a recent laptop or desktop with enough memory. Small models (a few billion parameters, quantised) run on many modern laptops; larger ones need a GPU with more memory or a machine with large unified memory (for example Apple Silicon). Check each model's memory requirements.
  • A runtime or app:
  • Ollama: command-line tool and background service that downloads and runs models, with a local API (including an OpenAI-compatible endpoint).
  • LM Studio: desktop app with a chat interface and model browser; can also serve a local API.
  • llama.cpp: the efficient engine underneath many tools, for advanced users.
  • A model sized for your hardware: start small and move up.

Hands-on 1: your first local model with Ollama

Install Ollama from its official website, then in a terminal:

# Download and chat with a model (choose one sized for your machine from the Ollama library)
ollama pull gpt-oss:20b
ollama run gpt-oss:20b "Summarise the pros and cons of running AI locally in 5 bullets."

# See what you have installed
ollama list

If a model is too slow or will not load, pick a smaller one or a more heavily quantised variant.

Hands-on 2: call your local model from code

Ollama exposes an OpenAI-compatible API on your machine, so the same client code works locally:

# pip install openai
from openai import OpenAI

local = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")  # key unused locally

resp = local.chat.completions.create(
    model="gpt-oss:20b",
    messages=[
        {"role": "system", "content": "You tag customer feedback. Reply with one word: praise, complaint, question or suggestion."},
        {"role": "user", "content": "The delivery was late again and nobody replied to my email."},
    ],
)
print(resp.choices[0].message.content)

This makes it easy to prototype privately, then decide whether a task needs a hosted model.

Good local use cases

Great locallyBetter with a hosted frontier model
Summarising private notes, journals, draftsComplex multi-step reasoning and strategy
Tagging and classifying sensitive textCurrent research (needs web search)
Coding help on private code (with a capable model)Long agentic tasks across many tools
Learning, experimenting, offline travelHighest-quality writing in many languages

Security and responsibility

  • Download only from official sources and reputable model hubs; verify you have the model you intended.
  • Keep the runtime updated and do not expose its local API to the internet unless you secure it.
  • Secure the device: full-disk encryption, screen lock and backups; local data is only as safe as the laptop.
  • Respect licences (previous lesson).
  • Understand quality limits: local models can hallucinate like any other; verify important outputs.

Worked example: a consultant on the road

A consultant in Riyadh travels frequently and handles confidential interview notes that her contract bans from third-party services. She runs a mid-size open-weight model locally with Ollama on an encrypted laptop, uses it to summarise interview notes and draft themes offline, and uses her company's approved hosted assistant only for non-confidential tasks such as desk research. She compares a sample of local summaries against her own reading to confirm quality.

Hands-on

If your computer allows, install Ollama or LM Studio, run a small model, and compare its answer with a hosted assistant on the same non-sensitive task. Note quality and speed differences.

Choosing a first model

Start with a small, well-known instruction-tuned model from the runtime's library, test it on three of your real tasks, and only move up in size if quality is clearly insufficient. Keep a note of the model name, size, quantisation and what it was good or bad at, so colleagues can reproduce your setup.

Pitfalls

  • Downloading models or apps from unofficial sources.
  • Choosing a model too large for your memory, then concluding "local AI is useless".
  • Exposing a local API to the network without authentication.
  • Assuming local means accurate.

How to measure success

You have a working local setup for at least one private workflow, you know which tasks it handles well, and sensitive material that must not leave your device never does.

Key takeaways

  • Local AI keeps data on your device and works offline, but models are usually smaller than frontier hosted ones.
  • You need adequate memory (and ideally a GPU), a local AI app or runtime, and a model sized for your hardware.
  • Great for private notes, drafting, tagging and learning; less suited to complex reasoning and current research.
  • Download only from official sources, secure your device, and respect licences.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What is the main privacy benefit of running a model locally?
  2. A small local model gives weak answers on a complex analysis. What's the most reasonable response?
  3. Where should you download local AI apps and models from?

Put it into practice

If your computer allows, install a reputable local AI app, run a small model and compare its answer with a hosted assistant on a non-sensitive task. Note differences in quality and speed.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.