AI Free course · Certificate included

Open-Weight and Local AI: Run, Choose and Deploy Your Own Models

Licenses, model families, quantization, VRAM sizing, Ollama, llama.cpp, vLLM, edge AI, privacy and hybrid routing

  • Advanced
  • 6 h 59 min
  • 16 lessons in 6 modules
  • 2 h 15 min of video lectures
  • Updated Sep 2026
Preview lesson 1
Earn the badgeCertified Open-Weight AI Engineer

About this course

Open-weight models now handle a large share of real business work on hardware you control. This advanced, hands-on course teaches you to choose, run and deploy them responsibly. You will read licenses the way legal will (Apache-2.0, MIT, the Llama community license), shortlist current model families on evidence, and understand quantization formats such as GGUF, AWQ, GPTQ, FP8 and MXFP4. You will size memory with a real formula for weights and KV cache, choose hardware on 24-month cost, and run models with Ollama, llama.cpp and LM Studio, including structured output and tool calling. You will serve at scale with vLLM, compare engines like SGLang fairly, put small models on devices, secure the model supply chain, respect data residency, and benchmark honestly. Finally you will build a hybrid router and a private, cited, permission-aware RAG assistant with an evaluation report.

Tools you’ll use

  • Ollama
  • llama.cpp
  • LM Studio
  • vLLM
  • SGLang
  • MLX
  • Hugging Face Hub
  • OpenAI Python SDK
  • Pydantic
  • NumPy
  • LiteLLM
  • promptfoo

Skills

  • Open-weight LLMs
  • Model licensing
  • Quantization
  • GPU sizing
  • LLM serving
  • Edge AI
  • Private RAG
  • LLM evaluation

What you’ll be able to do

  1. Explain open-weight vs closed models and triage licenses per release
  2. Shortlist model families and pick the smallest model that passes your evaluation
  3. Estimate memory for weights and KV cache and choose quantization and hardware
  4. Run, package and secure local models with Ollama, llama.cpp and LM Studio
  5. Serve models at scale with vLLM and benchmark engines fairly
  6. Secure the model supply chain and design for data residency
  7. Build a hybrid local/API router and a private RAG assistant with an eval report

Curriculum

Syllabus

Modules
6
Lessons
16
Reading time
4 h
Assessment questions
25
  1. What open-weight means, how licenses differ by release, and a repeatable way to shortlist model families for your tasks.

    1. Open-weight vs closed models: what "open" really meansVideo lecture, 9′14 min
    2. Licenses decoded: Apache-2.0, MIT, Llama and custom termsVideo lecture, 8′14 min
    3. Model families in 2026 and how to shortlistVideo lecture, 9′15 min
  2. Final assessment25 questions · 40 minutes · pass mark 80%

Your certificate

Finish with a credential anyone can check

Earn the Certified Open-Weight AI Engineer badge: The holder can choose open-weight models on license and evidence, size memory and hardware with the weights and KV-cache formula, run and secure models locally with Ollama and llama.cpp, serve them at scale with vLLM, apply data-residency and supply-chain controls, benchmark honestly, route between local and API models, and ship a private, cited, permission-aware RAG assistant.

Completed all lessons and scored at least 80% on the final assessment.

  • A public verification page
  • A PDF certificate to download
  • An Open Badge you can share
  • One click to your LinkedIn profile

Final assessment

  • 25questions drawn from a larger pool
  • 40 mintime limit
  • 80%pass mark
  • 3attempts per 24 hours

Start learning today. It’s free.

Every lesson is free to read. A free account saves your progress, unlocks the final assessment and issues your certificate.