Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsThe open-weight landscape · Lesson 1 of 16
Open-weight vs closed models: what "open" really means
Video lecture
Open-weight vs closed models: what "open" really means
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Open-weight AI: what "open" really means
Imagine your company's most sensitive document. A contract, a patient file, a payroll sheet. Now imagine asking an AI to summarize it, and knowing for certain that not a single word ever leaves the room. That is the promise of open-weight models. In this course you will learn to choose them, run them, size the hardware, serve them at scale and build a private assistant. Let's start by getting the words right, because the words drive the decisions.
0:34 Analogy: taxi vs owning a car
Here is an analogy that makes the whole decision easier. Using a closed model through an API is like taking a taxi. You get a great driver, you pay per trip, and you never worry about fuel, servicing or insurance. But you cannot choose the route engine, the company can change its prices, and your conversation happens inside someone else's car. Running an open-weight model is like owning a car. It costs money up front, you handle the servicing, and it may not be the fastest car on the road. But it goes where you want, when you want, and nobody else is in the passenger seat. Most businesses end up with both: a car for daily trips and a taxi for the occasional long journey.
1:29 Three categories
There are three categories. Closed models are only available through a vendor's API or app. You never see the weights, and the vendor decides when versions change or retire. Open-weight models let you download the trained parameters, the actual numbers that make the model work, under a license. And open source AI, in the strict sense used by the Open Source Initiative, goes further: enough information about data and code that others could study and rebuild the system. Most of the famous "open" models, Llama, Gemma, Qwen, Mistral, DeepSeek, Phi and gpt-oss, are open-weight. Keep that distinction in your pocket. It matters when a lawyer or auditor asks what exactly you are using.
2:18 What you gain
So what do you actually gain? First, data control. Prompts and documents stay on your laptop, your server or your private cloud. Second, version stability. The file you tested is the file you run, with no silent update on a Tuesday afternoon. Third, a different cost shape. You pay for hardware and electricity instead of per token. Fourth, customization, from fine-tuning to quantization. And fifth, offline use, on a ship, in a factory or on a phone.
2:51 What you give up
Now the honest part. The strongest closed models still tend to lead on the hardest reasoning and long agentic tasks, though the gap moves every few months and depends heavily on the task. You also take on operations: uptime, scaling, patches, monitoring. The safety filters an API vendor wraps around its model are now your responsibility. And "free tokens" is a myth. Idle GPUs and engineering hours cost real money.
3:21 The four questions
Here is a simple frame I use with clients. Four questions. One, data: must any of this data stay inside our boundary? Two, quality: on our own test set, does an open model hit the bar, maybe with retrieval or fine-tuning? Three, volume and latency: is traffic steady and high enough that owning inference pays off? Four, capability: can we operate this safely? A yes on data pulls you toward open weights. A no on quality pulls you toward an API, at least for that task. Most teams land on a hybrid.
4:01 Worked example: Dubai agency (illustrative)
Let's run it on a real-style example. A real-estate agency in Dubai wants an assistant that summarizes tenancy contracts full of passport and Emirates ID numbers. Data: sensitive, clients expect it to stay in the country. Quality: on sixty redacted contracts, a mid-size open model with retrieval gets fifty-five right, a frontier API gets fifty-seven. These numbers are illustrative, but the pattern is common. Volume: about four hundred questions a day in office hours, easy for one workstation GPU. Operations: their IT partner already runs servers. Decision: self-host, and route unusual clause questions to a human.
4:43 Hands-on
Your hands-on step is tiny but important. Install Ollama, then run one small model from your terminal and ask it a question. Then list your installed models and look at the size. That file on your disk is the model. No account, no API key, no network once it is downloaded. The exact commands are in the lesson text. From here we will learn to read licenses, pick model families, squeeze models onto your hardware and serve them to a whole team.
5:19 Simple example: freelance translator
Let me give you a simple example first, before the business one. A freelance translator in Lahore gets client documents that are under a non-disclosure agreement. Each week she translates about forty pages and wants help drafting first versions. The data rule is strict: nothing may leave her machine. Quality: a mid-size open model gets her first drafts to a level she can edit quickly. Volume: small and steady. Operations: she only needs one laptop and a runner like Ollama. Four questions, four answers, and the decision is obvious. Local, on her laptop. Notice she did not need a benchmark leaderboard to decide. She needed her own documents and one afternoon of testing.
6:08 Three mistakes to avoid
Three mistakes to avoid. First, calling a model open source in a policy or a contract when it is really open-weight with use restrictions. Words in contracts matter. Second, testing an open model only on your very hardest task, deciding it is not good enough, and missing the eighty percent of routine work where it would be perfect. Test across your real task mix. Third, forgetting the operations bill. Someone has to patch the server, watch the dashboards and plan upgrades, and that person's time belongs in your cost comparison. If you avoid those three, you are already ahead of most teams.
6:52 Try this now
Try this now. Open a note and write down one task you did this week where you pasted company information into an AI tool. Ask yourself the four questions for that one task: does the data need to stay inside, does a smaller model reach the quality bar, is the volume steady, and could someone on your team operate it? If you answered yes to data and yes to capability, that task is your first candidate for a local model. Keep the note. By the end of module three you will have that task running locally.
7:34 Watch me do it
Watch me do it. I will take one real workload and run the four questions live. The workload: summarizing supplier contracts for a procurement team in Karachi. First, data. I open a sample contract and highlight what is sensitive: bank details, pricing, a named contact. Verdict: it should stay inside. Second, quality. I pull ten redacted contracts, run them through a local eight billion model and a cloud model, and score each summary against a three-point checklist: parties, value, renewal date. The local model gets nine of ten fully right, the cloud model ten. Third, volume: about sixty contracts a day, steady. Fourth, capability: the IT team already runs a Linux server. I write the decision in one line on the note: local, with a human check on renewal dates, revisit in three months. That took me about forty minutes, and now the decision has evidence behind it instead of opinions.
8:40 Recap
Quick recap. Most open models are open-weight, not fully open source. They give you control, stability, customization and offline use, in exchange for some peak capability and a lot of operational responsibility. Decide with four questions: data, quality, volume, capability. Your next step: list three AI workloads in your organization and label each one local, API or hybrid. We will come back to that list at the end of the course.
Why this matters now
In 2026 you can download models that handle drafting, coding, retrieval-augmented question answering and tool calling well enough for a large share of business workloads, and run them on hardware you control. That changes three decisions every AI team makes: where data goes, what each task costs, and who controls the roadmap. This course teaches you to make those decisions on evidence, then build and operate the result.
Three terms people mix up
| Term | What you get | Example |
|---|---|---|
| Closed (API-only) | Access through a vendor's API or app. No weights. The vendor controls versions, pricing and retirement. | Frontier models from the major labs served via their APIs |
| Open-weight | The trained parameters are downloadable under a license. Training data and full training code usually are not. | Llama, Gemma, Qwen, Mistral, DeepSeek, Phi, gpt-oss |
| Open source AI (OSI definition) | Weights plus enough information about data and code to study and rebuild the system, under OSI-approved terms. | Fully open research models that publish data and recipes |
Most "open" LLMs you will use are open-weight, not open source in the strict sense. That distinction matters for audits ("can we explain what the model was trained on?") and for license obligations (covered in the next lesson).
What you gain with open weights
- Data control. Prompts and documents never leave your laptop, VPC or data center. For a Karachi clinic, a Riyadh law firm or a London HR team, that can be the difference between "not allowed" and "approved".
- Version stability. The file you tested is the file you run. No silent model updates, no forced migration when an API model is retired.
- Cost shape. You pay for hardware and energy (fixed-ish) rather than per token (variable). At high, steady volume this can be cheaper; at low or spiky volume it usually is not.
- Customization. You can fine-tune, quantize, prune, distil and change sampling in ways APIs do not expose.
- Offline and edge. Field teams, factories, aircraft, ships and phones can run models with no network.
What you give up
- Peak capability. The strongest closed models typically still lead on the hardest reasoning, long-horizon agentic work and multimodal tasks. The gap varies by task and changes month to month, so you measure it on your own tasks.
- Operations. You now own uptime, scaling, security patches, GPU procurement, monitoring and model upgrades.
- Built-in safety layers. API vendors wrap models in abuse monitoring and safety classifiers. With open weights, those guardrails are your job.
- Hidden costs. Engineering time, idle GPUs and evaluation work are real costs even if tokens are "free".
A decision frame: the four questions
- Data: Is there any data in this workload that must not leave our boundary (regulated, contractual, or simply sensitive)?
- Quality bar: On our own evaluation set, does an open model reach the bar, possibly with RAG or fine-tuning?
- Volume and latency: Is traffic high and steady enough, or latency-sensitive enough, that owning inference pays off?
- Capability to operate: Do we have (or can we buy) the skills to run it safely?
If the answer to 1 is "yes", open weights (or a private deployment of a closed model through your cloud provider) move to the top of the list. If 2 is "no", use a stronger API model for that task and revisit later. Many teams end up hybrid: local or self-hosted models for sensitive or high-volume tasks, and API models for the hardest cases. Module 6 builds that router.
Worked example: a Dubai real-estate agency
The agency wants an assistant that summarizes tenancy contracts and answers staff questions about them. Contracts contain passport numbers and Emirates ID data.
- Data: sensitive personal data; the agency's clients expect it to stay in the UAE. Open-weight is attractive.
- Quality: a test on 60 real (redacted) contracts shows a mid-size open model with retrieval answers 55 of 60 questions correctly; a frontier API model answers 57. The two misses are about unusual clauses that staff escalate anyway.
- Volume: around 400 queries a day, steady during office hours. A single workstation GPU handles this comfortably.
- Operations: their IT partner already runs on-premise servers.
Decision: self-host an open-weight model with retrieval, route clause-interpretation questions flagged as "unusual" to a human. Revisit quarterly. (Numbers are illustrative.)
Hands-on: your first local model in two minutes
Install Ollama from its official site, then in a terminal:
# Pull and chat with a small open-weight model (tags change; check the Ollama library)
ollama run qwen3:8b "In two sentences, what is an open-weight model?"
# See what is installed and how big each model is on disk
ollama listNotice the download size: that file is the model. Everything else in this course builds on that fact.
Pitfalls
- Calling a model "open source" in a contract or policy when it is open-weight with use restrictions.
- Comparing an open model on your hardest task only; test across your real task mix.
- Forgetting the operations bill: people, monitoring, upgrades.
How to measure success
You can state, for each AI workload you own, which of the four questions pushes it toward open weights, closed APIs or hybrid, and you have a small evaluation set to prove the quality claim.
Key takeaways
- Most "open" LLMs are open-weight: downloadable parameters under a license, not full open source
- Open weights give data control, version stability, customization and offline use
- You trade away some peak capability and take on operations and safety work
- Decide with four questions: data, quality bar, volume/latency, ability to operate
- Hybrid (local plus API) is the common end state
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
List three AI workloads in your organization. For each, answer the four decision questions and mark it local, API or hybrid. Keep this list; you will revisit it in Module 6.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.