---
title: "Open-Weight and Local AI: Run, Choose and Deploy Your Own…"
description: "Licenses, model families, quantization, VRAM sizing, Ollama, llama.cpp, vLLM, edge AI, privacy and hybrid routing"
url: https://optimizeall.com/learn/open-source-and-local-llms
updated: 2026-10-05
---

AI · Advanced · 419 minutes · free · updated Sep 2026

# Open-Weight and Local AI: Run, Choose and Deploy Your Own Models

Licenses, model families, quantization, VRAM sizing, Ollama, llama.cpp, vLLM, edge AI, privacy and hybrid routing

- **Lessons:** 16 in 6 modules
- **Video lectures:** 16 lectures, 135 minutes
- **Updated:** Sep 2026

## Tools you'll use

- Ollama
- llama.cpp
- LM Studio
- vLLM
- SGLang
- MLX
- Hugging Face Hub
- OpenAI Python SDK
- Pydantic
- NumPy
- LiteLLM
- promptfoo

[Start the course](https://optimizeall.com/learn/open-source-and-local-llms/open-weight-vs-closed-models)

## About this course

Open-weight models now handle a large share of real business work on hardware you control. This advanced, hands-on course teaches you to choose, run and deploy them responsibly. You will read licenses the way legal will (Apache-2.0, MIT, the Llama community license), shortlist current model families on evidence, and understand quantization formats such as GGUF, AWQ, GPTQ, FP8 and MXFP4. You will size memory with a real formula for weights and KV cache, choose hardware on 24-month cost, and run models with Ollama, llama.cpp and LM Studio, including structured output and tool calling. You will serve at scale with vLLM, compare engines like SGLang fairly, put small models on devices, secure the model supply chain, respect data residency, and benchmark honestly. Finally you will build a hybrid router and a private, cited, permission-aware RAG assistant with an evaluation report.

## What you will learn

- Explain open-weight vs closed models and triage licenses per release
- Shortlist model families and pick the smallest model that passes your evaluation
- Estimate memory for weights and KV cache and choose quantization and hardware
- Run, package and secure local models with Ollama, llama.cpp and LM Studio
- Serve models at scale with vLLM and benchmark engines fairly
- Secure the model supply chain and design for data residency
- Build a hybrid local/API router and a private RAG assistant with an eval report

## Before you start

- [Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)

## Course content

### The open-weight landscape

What open-weight means, how licenses differ by release, and a repeatable way to shortlist model families for your tasks.

- [Open-weight vs closed models: what "open" really means](https://optimizeall.com/learn/open-source-and-local-llms/open-weight-vs-closed-models): 14 min
- [Licenses decoded: Apache-2.0, MIT, Llama and custom terms](https://optimizeall.com/learn/open-source-and-local-llms/licenses-decoded): 14 min
- [Model families in 2026 and how to shortlist](https://optimizeall.com/learn/open-source-and-local-llms/model-families-and-selection): 15 min

### Quantization and hardware sizing

How quantization works and which formats to use, the memory formula for weights and KV cache, and how to choose and cost hardware.

- [Quantization: GGUF, AWQ, GPTQ, FP8 and MXFP4](https://optimizeall.com/learn/open-source-and-local-llms/how-quantization-works): 15 min
- [VRAM and memory sizing: the formula](https://optimizeall.com/learn/open-source-and-local-llms/vram-sizing-math): 15 min
- [Hardware choices: laptops, workstations, servers and cloud GPUs](https://optimizeall.com/learn/open-source-and-local-llms/hardware-choices): 14 min

### Running models locally

Hands-on with Ollama, llama.cpp and LM Studio, then turning a local model into a reliable component with schemas and tool calling.

- [Ollama hands-on: models, Modelfiles and the API](https://optimizeall.com/learn/open-source-and-local-llms/ollama-hands-on): 16 min
- [llama.cpp and LM Studio: control and convenience](https://optimizeall.com/learn/open-source-and-local-llms/llama-cpp-and-lm-studio): 15 min
- [Local models for real work: structured output, tools and reliability](https://optimizeall.com/learn/open-source-and-local-llms/local-tools-and-structured-output): 15 min

### Serving at scale

Production serving with vLLM, choosing between SGLang, TensorRT-LLM, llama.cpp and others, and measuring latency and throughput properly.

- [Serving at scale with vLLM](https://optimizeall.com/learn/open-source-and-local-llms/vllm-production-serving): 16 min
- [Choosing a serving engine: SGLang, TGI, TensorRT-LLM and friends](https://optimizeall.com/learn/open-source-and-local-llms/choosing-a-serving-engine): 14 min

### Edge AI, privacy and honest benchmarking

Small models on devices, data residency and model supply-chain security, and evaluation practices that produce decisions you can defend.

- [On-device AI and small language models](https://optimizeall.com/learn/open-source-and-local-llms/on-device-and-small-models): 15 min
- [Privacy, data residency and securing self-hosted models](https://optimizeall.com/learn/open-source-and-local-llms/privacy-residency-and-security): 15 min
- [Benchmarking responsibly: leaderboards, evals and honest comparisons](https://optimizeall.com/learn/open-source-and-local-llms/benchmarking-responsibly): 15 min

### Hybrid routing and capstone

Route requests between local and API models by sensitivity, task and availability, then build and evaluate a private RAG assistant.

- [Hybrid routing between local and API models](https://optimizeall.com/learn/open-source-and-local-llms/hybrid-routing-local-and-api): 16 min
- [Capstone: a private RAG assistant on a laptop or server](https://optimizeall.com/learn/open-source-and-local-llms/capstone-private-rag-assistant): 20 min

## Certificate: Certified Open-Weight AI Engineer

The holder can choose open-weight models on license and evidence, size memory and hardware with the weights and KV-cache formula, run and secure models locally with Ollama and llama.cpp, serve them at scale with vLLM, apply data-residency and supply-chain controls, benchmark honestly, route between local and API models, and ship a private, cited, permission-aware RAG assistant.

- **Final assessment:** 25 questions, 40 minutes
- **Passing score:** 80%
