Gemini, Microsoft Copilot, Perplexity & the AI Tool LandscapeOpen-weight models and local AI · Lesson 13 of 19

Open-weight models explained

Article · 16 min · 8 min lecture

Video lecture

Open-weight models explained

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Open-weight models

  • What “open” really means
  • When they beat closed models
  • Licences and hardware

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Closed, open-weight and open source

  • Closed (proprietary) models are accessed through a vendor's app or API; you never see the weights. Examples: the frontier models behind ChatGPT, Claude and Gemini.
  • Open-weight models publish their trained weights for download under a licence, so you can run them on your own hardware or cloud, fine-tune them, and control where data goes. Examples in 2026 include Meta's Llama family, Mistral models, Alibaba's Qwen, DeepSeek, Google's Gemma, and OpenAI's gpt-oss models (released in 2025).
  • Open source in the strict sense (as defined by the Open Source Initiative's Open Source AI Definition) also requires sufficient information about training data and code to study and modify the system. Many "open" models are open-weight but not fully open source.

Why businesses choose open-weight models

BenefitWhat it means in practice
Control and privacyData can stay on your servers or in your chosen cloud region
CustomisationFine-tune on your domain data or style; build narrow, efficient models
Cost at scaleFor high-volume, narrow tasks, self-hosting can be cheaper than per-token API pricing (depends heavily on utilisation)
IndependenceLess exposure to one vendor's price or policy changes
Offline or air-gappedPossible for sensitive environments

The trade-offs

  • Capability gap: the strongest closed frontier models usually still lead on the hardest reasoning and agentic tasks, although the gap varies by task and narrows with each release. Test on your tasks.
  • Operational burden: hosting, scaling, monitoring, security patching and evaluation are your job.
  • Safety responsibility: you must add guardrails, content filtering and abuse monitoring that a vendor would otherwise provide.
  • Licence obligations: terms vary widely.

Licences: read them

Open-weight licences range from permissive (for example Apache 2.0 or MIT, which several Mistral, Qwen, DeepSeek and gpt-oss models use) to custom licences with conditions (such as Meta's Llama community licence and Google's Gemma terms), which may include acceptable-use policies, attribution requirements or restrictions for very large services. Before commercial use:

  1. Find the official licence on the model's page (vendor site or Hugging Face model card).
  2. Check commercial use, attribution, acceptable-use policy, and any user-count or competitor restrictions.
  3. Check whether derivatives (fine-tuned versions) inherit the licence.
  4. Record the licence version with the model version you deploy.

Sizes, quantisation and hardware

  • Models come in sizes measured in parameters (for example, a few billion up to hundreds of billions). Bigger is usually more capable and needs more memory.
  • Mixture-of-experts models activate only part of the network per token, improving speed for their size.
  • Quantisation compresses weights (for example to 8-bit or 4-bit) so models fit in less memory, with some quality loss. Formats such as GGUF are common for local use.
  • Rough rule: the model file must fit in your GPU memory (or unified memory on Apple Silicon) with room to spare for context.

Where to find and run them

  • Hugging Face hosts most open-weight models with model cards, licences and evaluation notes.
  • Local runtimes: Ollama, LM Studio, llama.cpp (next lesson).
  • Cloud: major clouds and specialist providers host open-weight models behind APIs, often OpenAI-compatible, so you get the control benefits without owning hardware.

Worked example: a law firm's document tagging

A mid-size law firm in London must classify thousands of internal documents but its client contracts forbid sending them to external AI providers. It deploys an open-weight model on its own cloud tenancy, fine-tunes a small model on 2,000 labelled examples, and evaluates it against a held-out set of 300 documents. The small tuned model meets the accuracy bar for tagging; complex legal analysis still goes to lawyers. The licence (permissive) and data flow are documented for the firm's risk committee.

Hands-on

Pick one open-weight model family. Find its official licence and summarise in three bullets whether and how it permits commercial use, plus any conditions.

Evaluating an open-weight model for your task

Use the same test-set approach as for any tool (Lesson 7.1): 20 to 50 real examples with expected outputs, scored on accuracy, instruction-following and honesty. Compare the open-weight candidate with a hosted model on identical inputs. Include the total cost of running it (hardware or cloud hours, engineering time, monitoring), not just the absence of per-token fees.

Pitfalls

  • Assuming "open" means "no conditions".
  • Choosing the biggest model your hardware can barely run.
  • Skipping evaluation because "it's our own model".
  • Forgetting guardrails and security on self-hosted endpoints.

How to measure success

You can explain for any model you use: its licence terms, where it runs, how it was evaluated on your tasks, and what guardrails surround it.

Key takeaways

  • Closed models are accessed via apps or APIs; open-weight models publish weights under a licence; strict open source also releases enough data and code to study and modify.
  • Open-weight models (Llama, Mistral, Qwen, DeepSeek, Gemma, gpt-oss) offer control, privacy, customisation and cost at scale, at the price of operations and safety work.
  • Always read the licence: permissive (Apache 2.0, MIT) vs custom terms with conditions; record licence and model versions.
  • Choose per job and size to your hardware; quantisation trades some quality for memory.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What does 'open-weight' mean?
  2. An agency's client forbids sending data to external AI providers. Which option fits best?
  3. What is quantisation?

Put it into practice

Pick one open-weight model family. Find its official licence page and summarise in three bullets whether and how it permits commercial use.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.