Gemini, Microsoft Copilot, Perplexity & the AI Tool LandscapeOpen-weight models and local AI · Lesson 14 of 19
Running AI locally: a practical introduction
Video lecture
Running AI locally: a practical introduction
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Running AI locally
What if your AI assistant never sent a single word to anyone's servers? No subscription, no internet required, and your private notes never leave your laptop. That is what running AI locally offers. In this lecture you will learn what you need, set up your first local model, call it from code, and learn which jobs it handles well and which it does not.
0:28 Why local
Why run AI locally? Four reasons. Your prompts and documents stay on your device, which matters for private notes and confidential work. It works offline, on a plane or in a place with poor internet. There is no per use cost once it is set up. And it is one of the best ways to understand how these models actually work. The trade off is that local models are usually smaller, so they are weaker at complex reasoning, and they cannot search the web on their own.
1:06 The analogy
Think of local AI like a good home coffee machine. It is always available, nobody sees your order, and every cup is cheap once you own it. But it will not make every drink a professional barista can. For your daily espresso, it is perfect. For the complicated order, you still go to the café. Local and hosted models work together in exactly the same way.
1:35 What you need
You need three things. Hardware with enough memory. Small, quantised models run on many modern laptops, while bigger ones need a graphics card with more memory, or a machine with large unified memory, like recent Apple Silicon Macs. A runtime or app. Ollama is a command line tool with a background service and a local API. LM Studio is a desktop app with a chat window and model browser. llama dot c p p is the efficient engine underneath many tools. And a model sized for your hardware. Start small and move up.
2:15 Hands-on: Ollama
Let's set one up. Install Ollama from its official website. In a terminal, run ollama pull followed by the model name, choosing one sized for your machine from the Ollama library. Then ollama run with the same name and a question in quotes, and the answer streams back, generated entirely on your computer. Ollama list shows what you have installed. If a model is painfully slow or will not load, it is too big for your hardware, so pick a smaller or more heavily quantised version.
2:52 Simple example
A simple example. You are on a flight with your laptop in aeroplane mode. You paste your messy notes from yesterday's workshop into your local model and ask for five bullet points and a list of actions. The summary appears in seconds, and not a single byte left your laptop. For everyday private summarising, that is often all you need.
3:18 Hands-on: call it from code
You can also call your local model from code. Ollama exposes an OpenAI compatible API on your machine, at localhost port one one four three four. So you create the standard OpenAI client, point the base URL at your local address, and pass any placeholder key. Then send a system message, for example, tag customer feedback with one word, and a user message. The answer comes back from your own computer. This makes it easy to prototype privately, then decide whether a task truly needs a hosted model.
3:56 Great locally vs better hosted
Which jobs suit local models? Summarising private notes and drafts. Tagging and classifying sensitive text. Coding help on private code, with a capable model. Learning and experimenting, and working offline. Which are better hosted? Complex, multi step reasoning and strategy. Current research, which needs web search. Long agent tasks across many tools. And the very highest quality writing, especially across many languages.
4:23 Security and responsibility
A few security habits. Download apps and models only from official sources and reputable hubs. Keep the runtime updated. Do not expose its local API to the internet unless you secure it properly. Encrypt your disk, lock your screen and keep backups, because local data is only as safe as the laptop. Respect model licences. And remember that local models hallucinate like any other, so verify important outputs.
4:53 Business example: a consultant in Riyadh
A realistic business example. A consultant based in Riyadh travels constantly and handles confidential interview notes, which her contract bans from third party services. She runs a mid size open weight model locally with Ollama on an encrypted laptop, and uses it to summarise interviews and draft themes offline. Her company's approved hosted assistant is used only for non confidential tasks like desk research. She compares a sample of local summaries with her own reading to confirm the quality is good enough. She also documented her setup for colleagues with similar contracts. The runtime, the model name and size, the quantisation, and the tasks it handled well, interview themes and first draft summaries, plus the tasks it did not, complex strategic synthesis. Two colleagues copied the setup within a week, and the firm added it to its approved tools list for contract restricted material.
5:56 Common mistakes
Four common mistakes. Downloading models or apps from unofficial sources. Choosing a model far too large for your memory, and then concluding local AI is useless. Exposing a local API to the network without authentication. And assuming that because the model is local, it must be accurate. Private is not the same as correct.
6:19 Watch me do it, part 1
Let me set up the Riyadh consultant's laptop. With Ollama installed from the official site, I run ollama pull with a mid size model. It downloads with a progress bar. Then ollama run with the same name and a question, summarise the pros and cons of running AI locally in five bullets. The answer streams back, generated on the laptop. Out of curiosity I try a much larger model. It technically loads, but each word takes a second. Too big for this machine. Ollama list shows both, and I remove the large one.
6:59 Watch me do it, part 2
Now from code. I open the Python snippet, where the standard OpenAI client points its base URL at localhost port one one four three four, with a placeholder key. I switch off wifi to prove the point and run it. The feedback line comes back tagged complaint, entirely offline. Then the real job. I paste five confidential interview notes and ask for themes. The summaries are good enough for a first pass. Finally, I give the same public, non sensitive task to our approved hosted assistant and compare. Hosted is sharper, local is private. Each gets its job.
7:42 Recap and try this now
Recap. Local AI keeps data on your device, works offline and costs nothing per use, but smaller models have limits. Choose a runtime, size the model to your hardware, secure your device, and verify outputs. Try this now. If your computer allows, install Ollama or LM Studio, run a small model, and give it the same non sensitive task as your usual hosted assistant. Compare quality and speed, and decide which private jobs you will move to local.
Why run AI locally?
Running a model on your own computer means prompts and documents stay on your device, it works offline, and there is no per-use cost after setup. It is ideal for private notes, drafting, summarising personal documents, tagging, coding help on private code, and learning how models work. The trade-off: local models are usually smaller than frontier hosted models, so they are weaker at complex reasoning and have no built-in web search.
What you need
- Hardware: a recent laptop or desktop with enough memory. Small models (a few billion parameters, quantised) run on many modern laptops; larger ones need a GPU with more memory or a machine with large unified memory (for example Apple Silicon). Check each model's memory requirements.
- A runtime or app:
- Ollama: command-line tool and background service that downloads and runs models, with a local API (including an OpenAI-compatible endpoint).
- LM Studio: desktop app with a chat interface and model browser; can also serve a local API.
- llama.cpp: the efficient engine underneath many tools, for advanced users.
- A model sized for your hardware: start small and move up.
Hands-on 1: your first local model with Ollama
Install Ollama from its official website, then in a terminal:
# Download and chat with a model (choose one sized for your machine from the Ollama library)
ollama pull gpt-oss:20b
ollama run gpt-oss:20b "Summarise the pros and cons of running AI locally in 5 bullets."
# See what you have installed
ollama listIf a model is too slow or will not load, pick a smaller one or a more heavily quantised variant.
Hands-on 2: call your local model from code
Ollama exposes an OpenAI-compatible API on your machine, so the same client code works locally:
# pip install openai
from openai import OpenAI
local = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama") # key unused locally
resp = local.chat.completions.create(
model="gpt-oss:20b",
messages=[
{"role": "system", "content": "You tag customer feedback. Reply with one word: praise, complaint, question or suggestion."},
{"role": "user", "content": "The delivery was late again and nobody replied to my email."},
],
)
print(resp.choices[0].message.content)This makes it easy to prototype privately, then decide whether a task needs a hosted model.
Good local use cases
| Great locally | Better with a hosted frontier model |
|---|---|
| Summarising private notes, journals, drafts | Complex multi-step reasoning and strategy |
| Tagging and classifying sensitive text | Current research (needs web search) |
| Coding help on private code (with a capable model) | Long agentic tasks across many tools |
| Learning, experimenting, offline travel | Highest-quality writing in many languages |
Security and responsibility
- Download only from official sources and reputable model hubs; verify you have the model you intended.
- Keep the runtime updated and do not expose its local API to the internet unless you secure it.
- Secure the device: full-disk encryption, screen lock and backups; local data is only as safe as the laptop.
- Respect licences (previous lesson).
- Understand quality limits: local models can hallucinate like any other; verify important outputs.
Worked example: a consultant on the road
A consultant in Riyadh travels frequently and handles confidential interview notes that her contract bans from third-party services. She runs a mid-size open-weight model locally with Ollama on an encrypted laptop, uses it to summarise interview notes and draft themes offline, and uses her company's approved hosted assistant only for non-confidential tasks such as desk research. She compares a sample of local summaries against her own reading to confirm quality.
Hands-on
If your computer allows, install Ollama or LM Studio, run a small model, and compare its answer with a hosted assistant on the same non-sensitive task. Note quality and speed differences.
Choosing a first model
Start with a small, well-known instruction-tuned model from the runtime's library, test it on three of your real tasks, and only move up in size if quality is clearly insufficient. Keep a note of the model name, size, quantisation and what it was good or bad at, so colleagues can reproduce your setup.
Pitfalls
- Downloading models or apps from unofficial sources.
- Choosing a model too large for your memory, then concluding "local AI is useless".
- Exposing a local API to the network without authentication.
- Assuming local means accurate.
How to measure success
You have a working local setup for at least one private workflow, you know which tasks it handles well, and sensitive material that must not leave your device never does.
Key takeaways
- Local AI keeps data on your device and works offline, but models are usually smaller than frontier hosted ones.
- You need adequate memory (and ideally a GPU), a local AI app or runtime, and a model sized for your hardware.
- Great for private notes, drafting, tagging and learning; less suited to complex reasoning and current research.
- Download only from official sources, secure your device, and respect licences.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
If your computer allows, install a reputable local AI app, run a small model and compare its answer with a hosted assistant on a non-sensitive task. Note differences in quality and speed.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.