AI Free course · Certificate included
Open-Weight and Local AI: Run, Choose and Deploy Your Own Models
Licenses, model families, quantization, VRAM sizing, Ollama, llama.cpp, vLLM, edge AI, privacy and hybrid routing
- Advanced
- 6 h 59 min
- 16 lessons in 6 modules
- 2 h 15 min of video lectures
- Updated Sep 2026
About this course
Open-weight models now handle a large share of real business work on hardware you control. This advanced, hands-on course teaches you to choose, run and deploy them responsibly. You will read licenses the way legal will (Apache-2.0, MIT, the Llama community license), shortlist current model families on evidence, and understand quantization formats such as GGUF, AWQ, GPTQ, FP8 and MXFP4. You will size memory with a real formula for weights and KV cache, choose hardware on 24-month cost, and run models with Ollama, llama.cpp and LM Studio, including structured output and tool calling. You will serve at scale with vLLM, compare engines like SGLang fairly, put small models on devices, secure the model supply chain, respect data residency, and benchmark honestly. Finally you will build a hybrid router and a private, cited, permission-aware RAG assistant with an evaluation report.
Tools you’ll use
- Ollama
- llama.cpp
- LM Studio
- vLLM
- SGLang
- MLX
- Hugging Face Hub
- OpenAI Python SDK
- Pydantic
- NumPy
- LiteLLM
- promptfoo
Skills
- Open-weight LLMs
- Model licensing
- Quantization
- GPU sizing
- LLM serving
- Edge AI
- Private RAG
- LLM evaluation
What you’ll be able to do
- Explain open-weight vs closed models and triage licenses per release
- Shortlist model families and pick the smallest model that passes your evaluation
- Estimate memory for weights and KV cache and choose quantization and hardware
- Run, package and secure local models with Ollama, llama.cpp and LM Studio
- Serve models at scale with vLLM and benchmark engines fairly
- Secure the model supply chain and design for data residency
- Build a hybrid local/API router and a private RAG assistant with an eval report
Curriculum
Syllabus
- Modules
- 6
- Lessons
- 16
- Reading time
- 4 h
- Assessment questions
- 25
What open-weight means, how licenses differ by release, and a repeatable way to shortlist model families for your tasks.
- Open-weight vs closed models: what "open" really meansVideo lecture, 9′14 min
- Licenses decoded: Apache-2.0, MIT, Llama and custom termsVideo lecture, 8′14 min
- Model families in 2026 and how to shortlistVideo lecture, 9′15 min
How quantization works and which formats to use, the memory formula for weights and KV cache, and how to choose and cost hardware.
- Quantization: GGUF, AWQ, GPTQ, FP8 and MXFP4Video lecture, 8′15 min
- VRAM and memory sizing: the formulaVideo lecture, 9′15 min
- Hardware choices: laptops, workstations, servers and cloud GPUsVideo lecture, 9′14 min
Hands-on with Ollama, llama.cpp and LM Studio, then turning a local model into a reliable component with schemas and tool calling.
- Ollama hands-on: models, Modelfiles and the APIVideo lecture, 9′16 min
- llama.cpp and LM Studio: control and convenienceVideo lecture, 9′15 min
- Local models for real work: structured output, tools and reliabilityVideo lecture, 8′15 min
Production serving with vLLM, choosing between SGLang, TensorRT-LLM, llama.cpp and others, and measuring latency and throughput properly.
- Serving at scale with vLLMVideo lecture, 9′16 min
- Choosing a serving engine: SGLang, TGI, TensorRT-LLM and friendsVideo lecture, 8′14 min
Small models on devices, data residency and model supply-chain security, and evaluation practices that produce decisions you can defend.
- On-device AI and small language modelsVideo lecture, 8′15 min
- Privacy, data residency and securing self-hosted modelsVideo lecture, 8′15 min
- Benchmarking responsibly: leaderboards, evals and honest comparisonsVideo lecture, 8′15 min
Route requests between local and API models by sensitivity, task and availability, then build and evaluate a private RAG assistant.
- Hybrid routing between local and API modelsVideo lecture, 8′16 min
- Capstone: a private RAG assistant on a laptop or serverVideo lecture, 8′20 min
- Final assessment
Your certificate
Finish with a credential anyone can check
Earn the Certified Open-Weight AI Engineer badge: The holder can choose open-weight models on license and evidence, size memory and hardware with the weights and KV-cache formula, run and secure models locally with Ollama and llama.cpp, serve them at scale with vLLM, apply data-residency and supply-chain controls, benchmark honestly, route between local and API models, and ship a private, cited, permission-aware RAG assistant.
Completed all lessons and scored at least 80% on the final assessment.
- A public verification page
- A PDF certificate to download
- An Open Badge you can share
- One click to your LinkedIn profile
Final assessment
- 25questions drawn from a larger pool
- 40 mintime limit
- 80%pass mark
- 3attempts per 24 hours
Start learning today. It’s free.
Every lesson is free to read. A free account saves your progress, unlocks the final assessment and issues your certificate.