Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsQuantization and hardware sizing · Lesson 4 of 16

Quantization: GGUF, AWQ, GPTQ, FP8 and MXFP4

Article · 15 min · 8 min lecture

Video lecture

Quantization: GGUF, AWQ, GPTQ, FP8 and MXFP4

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Quantization

  • Shrink the model, keep most of the quality
  • Decode GGUF, AWQ, GPTQ, FP8, MXFP4
  • Choose a level with evidence

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

The one-sentence idea

Quantization stores each weight with fewer bits. A model trained in 16-bit precision (BF16 or FP16) uses 2 bytes per parameter. Store it in 8 bits and it halves; in 4 bits it quarters. Smaller files load faster, fit on cheaper hardware, and usually generate faster, because token generation is mostly limited by how fast you can read weights from memory.

The cost is precision. Each weight is rounded to one of fewer possible values, so outputs drift slightly from the original. Good quantization methods keep that drift small, especially at 8 and 5–6 bits; at 4 bits quality is usually close for mid-size and larger models; below 4 bits quality drops faster, especially for small models.

How it works (without the maths overload)

Weights are grouped into small blocks (for example 32 or 64 values). For each block, the method stores a scale (and sometimes a zero-point or minimum) in higher precision, plus a low-bit integer per weight. At inference the runtime multiplies back: weight ≈ scale × integer. Because each block has its own scale, outliers in one block do not ruin the precision of others.

The main families you will meet:

Format / methodWhere you meet itHow it worksTypical use
GGUF (llama.cpp)Ollama, llama.cpp, LM StudioA single-file format holding weights, tokenizer and metadata; supports many quant types such as Q8_0, Q6_K, Q5_K_M, Q4_K_M and the lower-bit IQ types. "K-quants" use nested blocks with their own scales. An importance matrix (imatrix) computed on sample text can improve low-bit quality.Laptops, desktops, Apple silicon, CPU and mixed CPU/GPU
AWQ (Activation-aware Weight Quantization)vLLM, SGLang, TransformersUses sample activations to find the small share of weights that matter most and protects them via scaling before 4-bit quantization.GPU serving at 4-bit
GPTQvLLM, TransformersPost-training quantization layer by layer, using approximate second-order information to compensate rounding error.GPU serving at 4/8-bit
FP8vLLM, SGLang, TensorRT-LLM on recent data-center GPUs8-bit floating point; recent NVIDIA and AMD accelerators support it natively; near-lossless for many models.High-throughput production
MXFP4gpt-oss native weightsA 4-bit microscaling floating-point format with shared block scales; gpt-oss ships this way, which is why gpt-oss-120b fits in 80 GB.Models released natively in low precision
bitsandbytes NF4Training (QLoRA), quick experiments4-bit "normal float" designed for normally distributed weights.Fine-tuning on small GPUs

Choosing a quantization level

A practical default ladder for local GGUF use:

  • Q8_0: near-original quality, half the size of 16-bit. Use when memory allows.
  • Q6_K / Q5_K_M: very close to original for most tasks; a sweet spot on workstations.
  • Q4_K_M: the common default for laptops; good quality for 7B+ models.
  • Q3 and IQ2/IQ3: only when a bigger model would not otherwise fit, and only after testing. A larger model at 3–4 bits often beats a smaller model at 8 bits, but not always, so measure.

For GPU serving, FP8 (on supported hardware) or 4-bit AWQ/GPTQ are the usual choices; vLLM and SGLang load these directly.

Don't forget the KV cache

Weights are not the only memory user. The KV cache stores attention keys and values for every token in the context, for every concurrent request. Long contexts and many users can make the KV cache larger than the weights. Most runtimes can quantize it too (for example 8-bit KV cache), with a small quality trade-off. The next lesson gives you the formula.

Worked example: fitting a 32B model on a 24 GB card

An engineering consultancy in Riyadh wants a ~32B model on engineers' 24 GB workstations for technical document Q&A.

  • 16-bit: about 64 GB for weights alone. Does not fit.
  • 8-bit: about 32 GB. Does not fit.
  • 4-bit (Q4_K_M effectively ~4.5–5 bits per weight): roughly 18–20 GB. Fits, leaving a few GB for KV cache at moderate context.

They compare the 32B model at Q4_K_M against a 14B model at Q8_0 on 50 of their own questions. In their (illustrative) test the 32B Q4 answers more technical questions correctly but is slower; they choose it for the Q&A tool and use the 14B for quick drafting.

Hands-on: make your own GGUF quantization

Most people download ready-made GGUF files, but doing it once teaches you what is inside. With a local clone of llama.cpp built per its README:

# 1) Convert a Hugging Face model directory (safetensors) to a 16-bit GGUF
python convert_hf_to_gguf.py ./my-model-dir --outfile my-model-f16.gguf --outtype f16

# 2) Optional: compute an importance matrix on representative text for better low-bit quality
./build/bin/llama-imatrix -m my-model-f16.gguf -f calibration.txt -o imatrix.dat

# 3) Quantize to Q4_K_M (add --imatrix imatrix.dat if you made one)
./build/bin/llama-quantize --imatrix imatrix.dat my-model-f16.gguf my-model-Q4_K_M.gguf Q4_K_M

# 4) Measure perplexity on held-out text to compare quant levels
./build/bin/llama-perplexity -m my-model-Q4_K_M.gguf -f heldout.txt

Script names and flags evolve; check the llama.cpp README for your version. Use calibration text that resembles your real traffic (for example, a mix of English and Arabic support messages).

Pitfalls

  • Judging quality by perplexity alone. Also run your task evaluation.
  • Using very low-bit quants on small models (1–4B), where degradation is sharper.
  • Mixing up files: a GGUF is for llama.cpp-family runtimes; AWQ/GPTQ/FP8 checkpoints are for GPU servers like vLLM.
  • Forgetting that re-uploaded quantizations must come from a trusted source.

How to measure success

For your chosen model, you have a table of quant levels vs file size vs task score vs tokens per second on target hardware, and you picked the smallest level whose task score is within your tolerance of the best.

Key takeaways

  • Quantization stores weights in fewer bits: smaller, faster, slightly less precise
  • GGUF is the llama.cpp/Ollama/LM Studio format; AWQ, GPTQ and FP8 are common for GPU servers
  • Q4_K_M is a common laptop default; Q5/Q6/Q8 when memory allows
  • A bigger model at 4-bit often beats a smaller one at 8-bit, but measure
  • The KV cache also consumes memory and can be quantized

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Why does 4-bit quantization usually increase generation speed on the same hardware?
  2. You serve a model on a data-center GPU with vLLM. Which format is least appropriate?
  3. What does an importance matrix (imatrix) do in llama.cpp quantization?

Put it into practice

Download two quantization levels (for example Q4_K_M and Q8_0) of the same model, run 20 of your tasks on each, and record size, speed and score.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.