---
title: "VRAM and memory sizing: the formula | Optimize All Academy"
description: "Why size on paper first GPUs and high-memory machines are expensive and often back-ordered. A ten-minute calculation tells you whether a model will fit…"
url: https://optimizeall.com/learn/open-source-and-local-llms/vram-sizing-math
updated: 2026-10-05
---

Open-Weight and Local AI: Run, Choose and Deploy Your Own Models · Quantization and hardware sizing · lesson 5 of 16 · 15 min

# VRAM and memory sizing: the formula

## Why size on paper first

GPUs and high-memory machines are expensive and often back-ordered. A ten-minute calculation tells you whether a model will fit, how much context you can offer, and how many users one card can serve, before you spend money. All numbers below are **illustrative estimates**; confirm on your hardware with your runtime's own memory report.

## The formula

Total memory needed is roughly:

```text
total ≈ weights + KV cache + runtime overhead

weights  (bytes) ≈ parameters × bits_per_weight ÷ 8
KV cache (bytes) ≈ 2 × layers × kv_heads × head_dim × bytes_per_kv_value × context_tokens × concurrent_sequences
overhead          ≈ activations, CUDA/Metal context, buffers: often 1–3 GB, sometimes more (measure)
```

Notes:

- `bits_per_weight` is the **effective** rate. Q4_K_M is roughly 4.5–5 bits per weight once scales are included; MXFP4 is a little over 4.
- For **MoE** models, use **total** parameters for weights memory.
- `kv_heads` is the number of **key/value** heads, which with grouped-query attention (GQA) is smaller than the number of attention heads. Find `num_hidden_layers`, `num_key_value_heads` and `head_dim` (or `hidden_size ÷ num_attention_heads`) in the model's `config.json`.
- The leading **2** is for keys *and* values.
- `bytes_per_kv_value` is 2 for FP16/BF16, 1 for 8-bit KV cache.

Some architectures use sliding-window or compressed attention (for example multi-head latent attention) and need less KV memory than this formula suggests. Treat the formula as an upper-bound estimate unless you know the architecture.

## Worked example 1: an 8B dense model on a laptop

Assume a model config with 32 layers, 8 KV heads, head_dim 128 (a common 8B configuration). Illustrative:

- **Weights at Q4_K_M:** 8B × ~4.8 bits ÷ 8 ≈ 4.8 GB.
- **KV cache per token (FP16):** 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes = **128 KiB per token**.
- **At 8,192 tokens, one user:** 8,192 × 128 KiB = 1 GiB.
- **At 32,768 tokens:** 4 GiB.
- **Total at 32k context:** ~4.8 + 4 + ~1.5 overhead ≈ **10–11 GB**.

So a 16 GB Apple silicon laptop can run it at 8k context comfortably; 32k context is tight once the OS and apps take their share.

## Worked example 2: a team server

A Karachi fintech wants to serve a 32B dense model (assume 64 layers, 8 KV heads, head_dim 128) to 20 concurrent analysts with 8k context each, on one GPU.

- **Weights, FP8:** 32 GB. **Weights, AWQ 4-bit:** ~18 GB.
- **KV per token (FP16):** 2 × 64 × 8 × 128 × 2 = 262,144 bytes = 256 KiB.
- **KV for 20 × 8,192 tokens:** 163,840 tokens × 256 KiB = 40 GiB.
- **Total (FP8 weights):** 32 + 40 + ~4 ≈ **76 GB** → an 80 GB GPU at full load is too tight for comfort.
- **Total (4-bit weights, 8-bit KV):** 18 + 20 + ~4 ≈ **42 GB** → fits on a 48 GB workstation card or comfortably on 80 GB.

In practice, not every analyst uses 8k tokens at the same moment. Servers like vLLM allocate KV cache in pages on demand (PagedAttention), so real concurrency is usually higher than this worst case. Size for the worst case you care about, then load-test.

## Speed: the second half of sizing

For single-user generation, a useful ceiling is:

```text
tokens/second (decode, batch 1) ≲ memory_bandwidth ÷ bytes_read_per_token
bytes_read_per_token ≈ weights size (dense) or active-parameter weights (MoE)
```

Example (illustrative): a GPU with ~1,000 GB/s bandwidth running a 4.8 GB model has a theoretical ceiling around 200 tokens/s; real numbers are lower. This also explains why MoE models generate fast: they read only the **active** experts per token. Prompt processing ("prefill") is compute-bound instead, which is why long prompts take noticeable time before the first token.

## Hands-on: a sizing calculator

```python
# vram_calc.py: rough memory estimate for LLM inference (illustrative, verify on hardware)
import json, sys

def estimate(params_b, bits, layers, kv_heads, head_dim, ctx, users, kv_bytes=2, overhead_gb=2.0):
    weights_gb = params_b * 1e9 * bits / 8 / 1e9
    kv_per_token = 2 * layers * kv_heads * head_dim * kv_bytes
    kv_gb = kv_per_token * ctx * users / 1e9
    return weights_gb, kv_gb, weights_gb + kv_gb + overhead_gb

if __name__ == "__main__":
    # usage: python vram_calc.py config.json 8.0 4.8 8192 4
    cfg = json.load(open(sys.argv[1]))
    params_b, bits, ctx, users = float(sys.argv[2]), float(sys.argv[3]), int(sys.argv[4]), int(sys.argv[5])
    layers = cfg["num_hidden_layers"]
    kv_heads = cfg.get("num_key_value_heads", cfg["num_attention_heads"])
    head_dim = cfg.get("head_dim") or cfg["hidden_size"] // cfg["num_attention_heads"]
    w, kv, total = estimate(params_b, bits, layers, kv_heads, head_dim, ctx, users)
    print(f"weights ≈ {w:.1f} GB | KV cache ≈ {kv:.1f} GB | total ≈ {total:.1f} GB (+ verify)")
```

Download only the model's `config.json`, run the calculator for three scenarios (solo, team, peak), then confirm with the runtime's own report (for example `ollama ps`, llama.cpp's load log, or vLLM's startup log showing KV cache capacity).

## Pitfalls

- Using attention heads instead of **KV** heads (overstates KV cache for GQA models).
- Forgetting concurrency: KV cache multiplies by simultaneous sequences.
- Sizing to 100% of memory. Leave headroom; the OS, drivers and fragmentation need space.
- Assuming advertised maximum context is affordable. Offer the context your users need.

## How to measure success

Your estimate lands within roughly 10–20% of the runtime's reported usage, and your load test confirms the concurrency you planned for.

## Video lecture: VRAM and memory sizing: the formula

Lecture coming soon · 13 chapters · about 9 minutes. Read the full transcript below.

1. Memory sizing: the formula
2. Analogy: a restaurant
3. total ≈ weights + KV cache + overhead
4. Details that matter
5. Laptop example (illustrative)
6. Team server (illustrative)
7. Speed ceiling
8. Hands-on + pitfalls
9. Quick mental math (illustrative)
10. Presenting sizing
11. Try this now
12. Watch me do it (illustrative)
13. Recap

## Lecture transcript

### Memory sizing: the formula

A client once asked me to approve a purchase order for a GPU server. I asked one question: how many users, at what context length? Nobody knew. Ten minutes with a calculator showed the card they wanted could serve about a third of their team at the context they needed. In this lesson you will learn that ten-minute calculation: how much memory a model needs, how much context you can offer, how many users fit on one card, and roughly how fast it will talk.

### Analogy: a restaurant

Think of GPU memory as a restaurant. The weights are the kitchen: a fixed space you need no matter how many guests arrive. The KV cache is the tables: every guest needs a table, and a guest with a long meal, a long conversation, needs a bigger table. Overhead is the corridors and storage. You can have the best kitchen in town, but if you only have room for three tables, you can only serve three guests at a time. Sizing is simply checking that the kitchen, the tables and the corridors all fit in the building.

### total ≈ weights + KV cache + overhead

Here is the whole formula. Total memory is roughly weights plus KV cache plus overhead. Weights in bytes equal parameters times effective bits per weight, divided by eight. KV cache equals two, for keys and values, times layers, times key-value heads, times head dimension, times bytes per value, times tokens in context, times concurrent sequences. And overhead covers the runtime and buffers, often a gigabyte or three, but measure it. You will find layers, key-value heads and head dimension in the model's config dot json file.

### Details that matter

Two details trip people up. First, use key-value heads, not attention heads. Most modern models use grouped-query attention, where many attention heads share fewer key-value heads, and that makes the cache much smaller. Second, use the effective bits. Q four K M is roughly four and a half to five bits per weight once you count the scales. And for mixture-of-experts models, weights memory uses the total parameters, all experts, even though only a few run per token.

### Laptop example (illustrative)

Let's do the laptop example together, with illustrative numbers. An eight billion parameter model with thirty-two layers, eight key-value heads and a head dimension of one twenty-eight. At Q four K M the weights are about four point eight gigabytes. The KV cache per token is two times thirty-two times eight times one twenty-eight times two bytes. That is a hundred and twenty-eight kibibytes per token. At eight thousand tokens, one gibibyte. At thirty-two thousand, four. Add overhead, and a thirty-two thousand token context needs around ten to eleven gigabytes. Comfortable at eight k on a sixteen gigabyte laptop, tight at thirty-two k.

### Team server (illustrative)

Now a team server. A Karachi fintech wants a thirty-two billion model for twenty analysts, each with eight thousand tokens of context. With a model of sixty-four layers, the KV cache is about two hundred fifty-six kibibytes per token. Twenty users times eight thousand tokens is about one hundred sixty-four thousand tokens, around forty gibibytes of cache. With eight-bit weights of thirty-two gigabytes, you are near seventy-six gigabytes. Too tight for an eighty gigabyte card. Switch to four-bit weights and an eight-bit KV cache and you are around forty-two gigabytes. Now it fits comfortably. Real servers allocate the cache on demand, so real capacity is often better, but you size for the worst case you care about and then load-test.

### Speed ceiling

Sizing is half the story. The other half is speed. When generating for a single user, a good ceiling is memory bandwidth divided by the bytes read per token. A card with about a thousand gigabytes per second of bandwidth, reading a model of about five gigabytes, tops out around two hundred tokens a second in theory, and real numbers are lower. That is also why mixture-of-experts models feel fast: they read only the active experts. Reading your prompt, called prefill, is different. It is compute-bound, which is why a very long prompt takes a moment before the first word appears.

### Hands-on + pitfalls

The lesson includes a small Python calculator. Download just the model's config file, run it for three scenarios, solo, team and peak, and then confirm with what your runtime actually reports. Ollama's ps command, the llama dot cpp load log, and vLLM's startup log all tell you real memory use. Common mistakes: using attention heads instead of key-value heads, forgetting concurrency, planning to use a hundred percent of memory, and promising the maximum advertised context when users need a fraction of it.

### Quick mental math (illustrative)

One more quick calculation, a tiny one, so you can do it in your head. A three billion parameter model at eight bits needs about three gigabytes for weights. Suppose its KV cache is about a hundred kilobytes per token. For four thousand tokens and one user, that is about four hundred megabytes. Add a gigabyte of overhead. Total: about four and a half gigabytes. That fits on almost any modern laptop GPU or a recent phone-class device with enough memory. You just sized a model in thirty seconds.

### Presenting sizing

Here is how to present sizing to a decision-maker. One slide, three scenarios: solo use, normal team load and peak. For each, show weights, cache and overhead against the card's memory, with a clear green or red. Add the expected tokens per second and the maximum users at your chosen context. Then state the assumptions: model, quantization, context and concurrency. When someone asks what if we double the users, you can answer in a minute, because the formula is right there.

### Try this now

Try this now. Pick a model you use, find its config file on its model page, and write down three numbers: number of layers, key-value heads and head dimension. Multiply two times layers times key-value heads times head dimension times two bytes. That is your KV cache per token. Multiply by the context length you actually need. Compare it with the size of the weights. Most people are shocked at how big the cache gets at long context.

### Watch me do it (illustrative)

Watch me do it. I open the config file of a fourteen billion parameter model. I find forty-eight hidden layers, eight key-value heads and a head dimension of one twenty-eight. Weights first: at Q four K M, about fourteen billion times four point eight bits, divided by eight, which is about eight point four gigabytes. KV cache per token: two times forty-eight times eight times one twenty-eight times two bytes, which is one hundred ninety-six thousand six hundred and eight bytes, about a hundred ninety-two kibibytes. My team wants eight thousand tokens each, for four people at once: that is thirty-two thousand tokens, about six gigabytes of cache. Add two gigabytes of overhead: about sixteen and a half gigabytes. My card has twenty-four. It fits, with headroom. Then I start the server and compare with the reported memory: within ten percent. All numbers here are illustrative, but the method is exactly what you will use.

### Recap

Recap. Total memory is weights plus KV cache plus overhead. Weights scale with parameters and bits. The KV cache scales with layers, key-value heads, head size, context and concurrent users. Speed is bounded by bandwidth divided by bytes per token. Estimate on paper, then confirm on hardware. Your next step: run the calculator for a model you are considering and write down the maximum users and context one card can support.

## Key takeaways

- Total memory ≈ weights + KV cache + overhead
- Weights ≈ parameters × effective bits ÷ 8; use total parameters for MoE
- KV cache ≈ 2 × layers × KV heads × head dim × bytes × tokens × concurrent sequences
- Decode speed is roughly bounded by memory bandwidth ÷ bytes read per token
- Estimate on paper, then confirm with the runtime report and a load test

## Try it

Fetch a model's config.json, run the sizing calculator for solo, team and peak scenarios, then compare with your runtime's reported memory.

- [Previous: Quantization: GGUF, AWQ, GPTQ, FP8 and MXFP4](https://optimizeall.com/learn/open-source-and-local-llms/how-quantization-works)
- [Next: Hardware choices: laptops, workstations, servers and cloud GPUs](https://optimizeall.com/learn/open-source-and-local-llms/hardware-choices)
- [All lessons of Open-Weight and Local AI: Run, Choose and Deploy Your Own Models](https://optimizeall.com/learn/open-source-and-local-llms)
