Gemini, Microsoft Copilot, Perplexity & the AI Tool LandscapeOpen-weight models and local AI · Lesson 13 of 19
Open-weight models explained
Video lecture
Open-weight models explained
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Open-weight models
Some of the most interesting AI in twenty twenty six is not behind a subscription at all. You can download it, run it on your own servers, and customise it. In this lecture you will learn what open weight models are, when they beat closed models for a business, the licence traps to avoid, and how model size and hardware fit together.
0:27 Why this matters
Why does this matter? Three reasons come up again and again. Some data cannot leave your control, because of client contracts, regulation or sheer sensitivity. Some tasks run at such high volume that per token pricing adds up. And depending on a single vendor means their price or policy change becomes your problem. Open weight models give you another option for each of those situations.
0:55 Three kinds of model
Let's get the vocabulary right. Closed models are used through a vendor's app or API, and you never see the weights. Open weight models publish their trained weights for download under a licence. Examples include Meta's Llama, Mistral, Alibaba's Qwen, DeepSeek, Google's Gemma and OpenAI's g p t oss models. And open source in the strict sense, as defined by the Open Source Initiative, also requires enough information about training data and code to study and modify the system. Many open models are open weight, but not fully open source.
1:34 The analogy
Here is an analogy. A closed model is like eating at a great restaurant. Excellent food, no kitchen to run, but you eat what is on the menu and you pay per plate. An open weight model is like getting the recipe. You can cook it at home, adjust the spices and cook in bulk more cheaply. But you need a kitchen, you need to learn to cook it well, and you do the washing up. Neither is better in general. It depends on what you are serving and how often.
2:14 Benefits
The benefits are real. Control and privacy, because data can stay on your servers or in your chosen cloud region. Customisation, because you can fine tune on your own data and style. Cost at scale, because for high volume, narrow tasks, self hosting can be cheaper than per token pricing, depending on how busy your hardware is. Independence from a single vendor. And even offline or air gapped use for sensitive environments.
2:45 Trade-offs
And the trade offs. The strongest closed frontier models usually still lead on the hardest reasoning and agent tasks, although the gap varies and narrows with each release. You run the operations, hosting, scaling, monitoring and security patches. You own safety, the guardrails and filtering a vendor would otherwise provide. And you must follow the licence, which brings us to the part people skip.
3:13 Read the licence
Licences range from permissive, like Apache two point zero or MIT, used by several Mistral, Qwen, DeepSeek and g p t oss models, to custom licences with conditions, such as Meta's Llama community licence and Google's Gemma terms. These may include acceptable use policies, attribution requirements, or restrictions for very large services. Before commercial use, find the official licence, check commercial use, attribution, acceptable use and any user count limits, check whether fine tuned versions inherit it, and record the licence version with the model you deploy.
3:51 Simple example
A simple example. A writer wants to summarise her private journal entries into weekly themes, but she does not want them on anyone's servers. A small open weight model running on her laptop, offline, does the job well. Summaries of personal notes do not need a frontier model. She checks the licence, which is permissive, and she is done. Small model, right job, full privacy.
4:19 Sizes and hardware
Now sizes and hardware. Models are measured in parameters, from a few billion to hundreds of billions. Bigger is usually more capable and needs more memory. Mixture of experts models activate only part of the network for each token, which helps speed. Quantisation compresses the weights, for example to eight or four bit, so a model fits in less memory, with some quality loss. The rough rule is that the model must fit in your GPU memory, or unified memory on Apple Silicon, with room left for the context.
4:58 Business example: a London law firm
A realistic business example. A mid size law firm in London needs to classify thousands of internal documents, but its client contracts forbid sending them to external AI providers. It deploys an open weight model on its own cloud tenancy and fine tunes a small model on two thousand labelled examples. Tested on three hundred held out documents, the small model meets the accuracy bar for tagging. Complex legal analysis still goes to lawyers. The licence and data flow are documented for the risk committee. A year later the firm reviewed the decision. A newer open weight model had been released under the same permissive licence, and the firm reran its three hundred document evaluation. The newer model scored higher, so they upgraded, updated the register with the new version and licence, and repeated the risk committee sign off. The process made the upgrade routine rather than risky.
6:02 Common mistakes
Four common mistakes. Assuming open means no conditions. Choosing the biggest model your hardware can barely run, which is slow and fragile. Skipping evaluation because it is our own model. And forgetting guardrails and security on self hosted endpoints, which then sit on the internet with no protection at all.
6:24 Watch me do it, part 1
Let me check a licence the way the law firm did. I open the model's page on Hugging Face. The model card shows the parameter count, the available formats and quantised versions, and a licence field. I click the licence and read it. Commercial use, permitted. Attribution, required in the documentation. Acceptable use policy, a list of prohibited uses. And for fine tuned versions, the same licence applies. I record the model name, version, licence and date checked in our register, alongside the risk committee's approval.
7:01 Watch me do it, part 2
Next, size. Our server's GPU has limited memory. The eight bit version of the larger model does not fit with room for context, so I mark it red. The four bit version fits, and a smaller model fits easily. Then the evaluation plan. Three hundred held out documents, already labelled by paralegals, and an accuracy target agreed with the practice lead. After fine tuning, the small model meets the target on tagging. Our notes say clearly, suitable for document tagging, not for legal analysis, which stays with lawyers.
7:39 Recap and try this now
Recap. Know the difference between closed, open weight and open source. Choose open weight for control, customisation, cost at scale and independence, and accept the operational and safety work. Read the licence every time, and size the model to your hardware. Try this now. Pick one open weight model family, find its official licence page, and summarise in three bullets whether and how it permits commercial use, and any conditions you would need to follow.
Closed, open-weight and open source
- Closed (proprietary) models are accessed through a vendor's app or API; you never see the weights. Examples: the frontier models behind ChatGPT, Claude and Gemini.
- Open-weight models publish their trained weights for download under a licence, so you can run them on your own hardware or cloud, fine-tune them, and control where data goes. Examples in 2026 include Meta's Llama family, Mistral models, Alibaba's Qwen, DeepSeek, Google's Gemma, and OpenAI's gpt-oss models (released in 2025).
- Open source in the strict sense (as defined by the Open Source Initiative's Open Source AI Definition) also requires sufficient information about training data and code to study and modify the system. Many "open" models are open-weight but not fully open source.
Why businesses choose open-weight models
| Benefit | What it means in practice |
|---|---|
| Control and privacy | Data can stay on your servers or in your chosen cloud region |
| Customisation | Fine-tune on your domain data or style; build narrow, efficient models |
| Cost at scale | For high-volume, narrow tasks, self-hosting can be cheaper than per-token API pricing (depends heavily on utilisation) |
| Independence | Less exposure to one vendor's price or policy changes |
| Offline or air-gapped | Possible for sensitive environments |
The trade-offs
- Capability gap: the strongest closed frontier models usually still lead on the hardest reasoning and agentic tasks, although the gap varies by task and narrows with each release. Test on your tasks.
- Operational burden: hosting, scaling, monitoring, security patching and evaluation are your job.
- Safety responsibility: you must add guardrails, content filtering and abuse monitoring that a vendor would otherwise provide.
- Licence obligations: terms vary widely.
Licences: read them
Open-weight licences range from permissive (for example Apache 2.0 or MIT, which several Mistral, Qwen, DeepSeek and gpt-oss models use) to custom licences with conditions (such as Meta's Llama community licence and Google's Gemma terms), which may include acceptable-use policies, attribution requirements or restrictions for very large services. Before commercial use:
- Find the official licence on the model's page (vendor site or Hugging Face model card).
- Check commercial use, attribution, acceptable-use policy, and any user-count or competitor restrictions.
- Check whether derivatives (fine-tuned versions) inherit the licence.
- Record the licence version with the model version you deploy.
Sizes, quantisation and hardware
- Models come in sizes measured in parameters (for example, a few billion up to hundreds of billions). Bigger is usually more capable and needs more memory.
- Mixture-of-experts models activate only part of the network per token, improving speed for their size.
- Quantisation compresses weights (for example to 8-bit or 4-bit) so models fit in less memory, with some quality loss. Formats such as GGUF are common for local use.
- Rough rule: the model file must fit in your GPU memory (or unified memory on Apple Silicon) with room to spare for context.
Where to find and run them
- Hugging Face hosts most open-weight models with model cards, licences and evaluation notes.
- Local runtimes: Ollama, LM Studio, llama.cpp (next lesson).
- Cloud: major clouds and specialist providers host open-weight models behind APIs, often OpenAI-compatible, so you get the control benefits without owning hardware.
Worked example: a law firm's document tagging
A mid-size law firm in London must classify thousands of internal documents but its client contracts forbid sending them to external AI providers. It deploys an open-weight model on its own cloud tenancy, fine-tunes a small model on 2,000 labelled examples, and evaluates it against a held-out set of 300 documents. The small tuned model meets the accuracy bar for tagging; complex legal analysis still goes to lawyers. The licence (permissive) and data flow are documented for the firm's risk committee.
Hands-on
Pick one open-weight model family. Find its official licence and summarise in three bullets whether and how it permits commercial use, plus any conditions.
Evaluating an open-weight model for your task
Use the same test-set approach as for any tool (Lesson 7.1): 20 to 50 real examples with expected outputs, scored on accuracy, instruction-following and honesty. Compare the open-weight candidate with a hosted model on identical inputs. Include the total cost of running it (hardware or cloud hours, engineering time, monitoring), not just the absence of per-token fees.
Pitfalls
- Assuming "open" means "no conditions".
- Choosing the biggest model your hardware can barely run.
- Skipping evaluation because "it's our own model".
- Forgetting guardrails and security on self-hosted endpoints.
How to measure success
You can explain for any model you use: its licence terms, where it runs, how it was evaluated on your tasks, and what guardrails surround it.
Key takeaways
- Closed models are accessed via apps or APIs; open-weight models publish weights under a licence; strict open source also releases enough data and code to study and modify.
- Open-weight models (Llama, Mistral, Qwen, DeepSeek, Gemma, gpt-oss) offer control, privacy, customisation and cost at scale, at the price of operations and safety work.
- Always read the licence: permissive (Apache 2.0, MIT) vs custom terms with conditions; record licence and model versions.
- Choose per job and size to your hardware; quantisation trades some quality for memory.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Pick one open-weight model family. Find its official licence page and summarise in three bullets whether and how it permits commercial use.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.