---
title: "Quantization: GGUF, AWQ, GPTQ, FP8 and MXFP4"
description: "The one-sentence idea Quantization stores each weight with fewer bits. A model trained in 16-bit precision (BF16 or FP16) uses 2 bytes per parameter…"
url: https://optimizeall.com/learn/open-source-and-local-llms/how-quantization-works
updated: 2026-10-05
---

Open-Weight and Local AI: Run, Choose and Deploy Your Own Models · Quantization and hardware sizing · lesson 4 of 16 · 15 min

# Quantization: GGUF, AWQ, GPTQ, FP8 and MXFP4

## The one-sentence idea

**Quantization stores each weight with fewer bits.** A model trained in 16-bit precision (BF16 or FP16) uses 2 bytes per parameter. Store it in 8 bits and it halves; in 4 bits it quarters. Smaller files load faster, fit on cheaper hardware, and usually generate faster, because token generation is mostly limited by how fast you can read weights from memory.

The cost is **precision**. Each weight is rounded to one of fewer possible values, so outputs drift slightly from the original. Good quantization methods keep that drift small, especially at 8 and 5–6 bits; at 4 bits quality is usually close for mid-size and larger models; below 4 bits quality drops faster, especially for small models.

## How it works (without the maths overload)

Weights are grouped into small **blocks** (for example 32 or 64 values). For each block, the method stores a **scale** (and sometimes a zero-point or minimum) in higher precision, plus a low-bit integer per weight. At inference the runtime multiplies back: `weight ≈ scale × integer`. Because each block has its own scale, outliers in one block do not ruin the precision of others.

The main families you will meet:

| Format / method | Where you meet it | How it works | Typical use |
|---|---|---|---|
| **GGUF** (llama.cpp) | Ollama, llama.cpp, LM Studio | A single-file format holding weights, tokenizer and metadata; supports many quant types such as `Q8_0`, `Q6_K`, `Q5_K_M`, `Q4_K_M` and the lower-bit `IQ` types. "K-quants" use nested blocks with their own scales. An **importance matrix (imatrix)** computed on sample text can improve low-bit quality. | Laptops, desktops, Apple silicon, CPU and mixed CPU/GPU |
| **AWQ** (Activation-aware Weight Quantization) | vLLM, SGLang, Transformers | Uses sample activations to find the small share of weights that matter most and protects them via scaling before 4-bit quantization. | GPU serving at 4-bit |
| **GPTQ** | vLLM, Transformers | Post-training quantization layer by layer, using approximate second-order information to compensate rounding error. | GPU serving at 4/8-bit |
| **FP8** | vLLM, SGLang, TensorRT-LLM on recent data-center GPUs | 8-bit floating point; recent NVIDIA and AMD accelerators support it natively; near-lossless for many models. | High-throughput production |
| **MXFP4** | gpt-oss native weights | A 4-bit microscaling floating-point format with shared block scales; gpt-oss ships this way, which is why gpt-oss-120b fits in 80 GB. | Models released natively in low precision |
| **bitsandbytes NF4** | Training (QLoRA), quick experiments | 4-bit "normal float" designed for normally distributed weights. | Fine-tuning on small GPUs |

## Choosing a quantization level

A practical default ladder for local GGUF use:

- **Q8_0**: near-original quality, half the size of 16-bit. Use when memory allows.
- **Q6_K / Q5_K_M**: very close to original for most tasks; a sweet spot on workstations.
- **Q4_K_M**: the common default for laptops; good quality for 7B+ models.
- **Q3 and IQ2/IQ3**: only when a bigger model would not otherwise fit, and only after testing. A larger model at 3–4 bits often beats a smaller model at 8 bits, but not always, so measure.

For GPU serving, FP8 (on supported hardware) or 4-bit AWQ/GPTQ are the usual choices; vLLM and SGLang load these directly.

## Don't forget the KV cache

Weights are not the only memory user. The **KV cache** stores attention keys and values for every token in the context, for every concurrent request. Long contexts and many users can make the KV cache larger than the weights. Most runtimes can quantize it too (for example 8-bit KV cache), with a small quality trade-off. The next lesson gives you the formula.

## Worked example: fitting a 32B model on a 24 GB card

An engineering consultancy in Riyadh wants a ~32B model on engineers' 24 GB workstations for technical document Q&A.

- 16-bit: about 64 GB for weights alone. Does not fit.
- 8-bit: about 32 GB. Does not fit.
- 4-bit (Q4_K_M effectively ~4.5–5 bits per weight): roughly 18–20 GB. Fits, leaving a few GB for KV cache at moderate context.

They compare the 32B model at Q4_K_M against a 14B model at Q8_0 on 50 of their own questions. In their (illustrative) test the 32B Q4 answers more technical questions correctly but is slower; they choose it for the Q&A tool and use the 14B for quick drafting.

## Hands-on: make your own GGUF quantization

Most people download ready-made GGUF files, but doing it once teaches you what is inside. With a local clone of llama.cpp built per its README:

```bash
# 1) Convert a Hugging Face model directory (safetensors) to a 16-bit GGUF
python convert_hf_to_gguf.py ./my-model-dir --outfile my-model-f16.gguf --outtype f16

# 2) Optional: compute an importance matrix on representative text for better low-bit quality
./build/bin/llama-imatrix -m my-model-f16.gguf -f calibration.txt -o imatrix.dat

# 3) Quantize to Q4_K_M (add --imatrix imatrix.dat if you made one)
./build/bin/llama-quantize --imatrix imatrix.dat my-model-f16.gguf my-model-Q4_K_M.gguf Q4_K_M

# 4) Measure perplexity on held-out text to compare quant levels
./build/bin/llama-perplexity -m my-model-Q4_K_M.gguf -f heldout.txt
```

Script names and flags evolve; check the llama.cpp README for your version. Use calibration text that resembles your real traffic (for example, a mix of English and Arabic support messages).

## Pitfalls

- Judging quality by perplexity alone. Also run your task evaluation.
- Using very low-bit quants on small models (1–4B), where degradation is sharper.
- Mixing up files: a GGUF is for llama.cpp-family runtimes; AWQ/GPTQ/FP8 checkpoints are for GPU servers like vLLM.
- Forgetting that re-uploaded quantizations must come from a trusted source.

## How to measure success

For your chosen model, you have a table of quant levels vs file size vs task score vs tokens per second on target hardware, and you picked the smallest level whose task score is within your tolerance of the best.

## Video lecture: Quantization: GGUF, AWQ, GPTQ, FP8 and MXFP4

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

1. Quantization
2. Analogy: color palettes
3. How it works
4. Why smaller is faster
5. The formats
6. The local ladder
7. Worked example: Riyadh consultancy (illustrative)
8. Don't forget the KV cache
9. Simple example: 8B on a 16 GB laptop
10. Hands-on + pitfalls
11. Try this now
12. Watch me do it
13. Recap

## Lecture transcript

### Quantization

Here is a party trick that is actually an engineering superpower. Take a model file that weighs sixteen gigabytes and shrink it to about five, then watch it run faster, while answering almost as well. That is quantization. In this lesson you will learn what it really does, decode the alphabet soup of GGUF, AWQ, GPTQ, FP8 and MXFP4, and pick the right level for your hardware with evidence rather than guesswork.

### Analogy: color palettes

Here is the analogy I like best. Imagine a photo with millions of colors. Now save it with a smaller palette of, say, two hundred and fifty-six colors. From arm's length, it looks almost the same, and the file is a fraction of the size. Look very closely at a sunset gradient, and you might see some banding. Quantization does the same to a model's weights. Eight bits is a rich palette, almost indistinguishable. Four bits is a careful smaller palette, still very good. Two bits is a poster with a handful of colors: recognizable, but you lose detail.

### How it works

The idea is simple. Models are trained with each weight stored in sixteen bits, two bytes. Quantization stores each weight in fewer bits, eight, six, five or four. Here is how it keeps quality. Weights are grouped into small blocks, maybe thirty-two values each. For every block you keep one precise scale number, and for every weight a small integer. At runtime you multiply them back together. Because every block has its own scale, one unusual weight does not ruin the precision of the rest.

### Why smaller is faster

Why does smaller mean faster? When a model generates a token, it has to read almost all of its weights from memory. On most hardware, that memory read is the bottleneck, not the arithmetic. So if every weight takes a quarter of the bytes, each token needs a quarter of the reading. That is why a four-bit model on the same laptop often streams text noticeably faster than the sixteen-bit original.

### The formats

Now the formats. GGUF is the single-file format used by llama dot cpp, Ollama and LM Studio. You will see names like Q eight zero, Q six K, Q five K M and Q four K M, which are different quantization recipes. AWQ protects the small set of weights that matter most, measured on sample activations, before going to four bits. GPTQ quantizes layer by layer and compensates the rounding error. Both are popular for GPU servers like vLLM. FP8 is an eight-bit floating-point format that recent data-center GPUs run natively. And MXFP4 is a four-bit microscaling format, the one gpt-oss ships in, which is why the one-twenty B model fits on a single eighty gigabyte GPU.

### The local ladder

How do you choose? For local use, start with this ladder. Q eight zero is near-original at half the size. Q six K and Q five K M are the sweet spot if memory allows. Q four K M is the common laptop default and works well for models around seven billion parameters and up. Three bits and below are for emergencies: when a bigger model would not otherwise fit, and only after testing. And here is the surprising part. A bigger model at four bits often beats a smaller model at eight bits. Often, not always. So you measure.

### Worked example: Riyadh consultancy (illustrative)

Let's fit a real-style case. An engineering consultancy in Riyadh wants a thirty-two billion parameter model on twenty-four gigabyte workstations. At sixteen bits the weights alone are about sixty-four gigabytes. No chance. At eight bits, thirty-two. Still no. At Q four K M, roughly eighteen to twenty gigabytes, which fits with room for a working context. They test it against a fourteen billion model at eight bits on fifty of their own questions. In their illustrative test, the bigger four-bit model wins on hard technical questions, the smaller one is quicker, so they use each where it fits.

### Don't forget the KV cache

Remember the other memory user: the KV cache. It stores attention information for every token in the conversation, for every user at once. With long documents and many users, it can outgrow the weights. Most runtimes can quantize the KV cache too, often to eight bits, with a small quality trade-off. We will put exact numbers on this in the next lesson.

### Simple example: 8B on a 16 GB laptop

A simple example. You have a laptop with sixteen gigabytes of memory and you want to run an eight billion parameter model. At sixteen bits, the weights alone are about sixteen gigabytes, so it will not fit alongside your operating system. At eight bits, about eight gigabytes: it fits, but tightly. At Q four K M, roughly five gigabytes: it fits comfortably with room for a decent context. So you try Q four K M first, run twenty of your own prompts, and if the answers look as good as the eight-bit version, you keep it.

### Hands-on + pitfalls

Your hands-on option: make a quantization yourself with llama dot cpp. Convert a model to a sixteen-bit GGUF, optionally compute an importance matrix on text that looks like your real traffic, quantize to Q four K M, then measure perplexity and, more importantly, run your task evaluation. The commands are in the lesson. The pitfalls: trusting perplexity alone, pushing tiny models to very low bits, and mixing up which files belong to which runtime.

### Try this now

Try this now. In Ollama or LM Studio, look up one model you use and find which quantization you are actually running. Ollama's show command will tell you, and LM Studio displays it next to the download. Many people are surprised to learn they have been running a four-bit model all along. Then download one level higher, run the same five prompts on both, and write down whether you can tell the difference.

### Watch me do it

Watch me do it. I have one model in two quantizations, Q four K M and Q eight zero, both downloaded as GGUF files. First I check sizes on disk: roughly five gigabytes versus eight and a half. Then I run llama bench on both with the same settings, and write down prompt processing and generation speed. The four-bit version generates noticeably faster on my laptop. Now quality. I run the same twenty prompts from my own work through both, at temperature zero, and paste the outputs side by side in a sheet. I score each pair: same quality, four-bit worse, or four-bit better. Result: seventeen the same, two slightly worse, one oddly better. The two worse ones are long numeric tables, where small errors crept in. So my decision: Q four K M for everyday drafting, Q eight zero for the table-heavy report workflow. Size, speed and task quality, all measured in under an hour.

### Recap

Recap. Quantization stores weights in fewer bits, so models get smaller and usually faster, with a small precision cost. GGUF is for local runtimes; AWQ, GPTQ and FP8 are for GPU servers; MXFP4 is how gpt-oss ships. Q four K M is a sensible default, but a table of size, speed and task score is what makes the decision. Your next step: run two quant levels of one model on twenty of your tasks and fill in that table.

## Key takeaways

- Quantization stores weights in fewer bits: smaller, faster, slightly less precise
- GGUF is the llama.cpp/Ollama/LM Studio format; AWQ, GPTQ and FP8 are common for GPU servers
- Q4_K_M is a common laptop default; Q5/Q6/Q8 when memory allows
- A bigger model at 4-bit often beats a smaller one at 8-bit, but measure
- The KV cache also consumes memory and can be quantized

## Try it

Download two quantization levels (for example Q4_K_M and Q8_0) of the same model, run 20 of your tasks on each, and record size, speed and score.

- [Previous: Model families in 2026 and how to shortlist](https://optimizeall.com/learn/open-source-and-local-llms/model-families-and-selection)
- [Next: VRAM and memory sizing: the formula](https://optimizeall.com/learn/open-source-and-local-llms/vram-sizing-math)
- [All lessons of Open-Weight and Local AI: Run, Choose and Deploy Your Own Models](https://optimizeall.com/learn/open-source-and-local-llms)
