Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsQuantization and hardware sizing · Lesson 5 of 16

VRAM and memory sizing: the formula

Article · 15 min · 9 min lecture

Video lecture

VRAM and memory sizing: the formula

13 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 13

Memory sizing: the formula

  • Will it fit?
  • How much context, how many users?
  • How fast will it generate?

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why size on paper first

GPUs and high-memory machines are expensive and often back-ordered. A ten-minute calculation tells you whether a model will fit, how much context you can offer, and how many users one card can serve, before you spend money. All numbers below are illustrative estimates; confirm on your hardware with your runtime's own memory report.

The formula

Total memory needed is roughly:

total ≈ weights + KV cache + runtime overhead

weights  (bytes) ≈ parameters × bits_per_weight ÷ 8
KV cache (bytes) ≈ 2 × layers × kv_heads × head_dim × bytes_per_kv_value × context_tokens × concurrent_sequences
overhead          ≈ activations, CUDA/Metal context, buffers: often 1–3 GB, sometimes more (measure)

Notes:

  • bits_per_weight is the effective rate. Q4_K_M is roughly 4.5–5 bits per weight once scales are included; MXFP4 is a little over 4.
  • For MoE models, use total parameters for weights memory.
  • kv_heads is the number of key/value heads, which with grouped-query attention (GQA) is smaller than the number of attention heads. Find num_hidden_layers, num_key_value_heads and head_dim (or hidden_size ÷ num_attention_heads) in the model's config.json.
  • The leading 2 is for keys and values.
  • bytes_per_kv_value is 2 for FP16/BF16, 1 for 8-bit KV cache.

Some architectures use sliding-window or compressed attention (for example multi-head latent attention) and need less KV memory than this formula suggests. Treat the formula as an upper-bound estimate unless you know the architecture.

Worked example 1: an 8B dense model on a laptop

Assume a model config with 32 layers, 8 KV heads, head_dim 128 (a common 8B configuration). Illustrative:

  • Weights at Q4_K_M: 8B × ~4.8 bits ÷ 8 ≈ 4.8 GB.
  • KV cache per token (FP16): 2 × 32 × 8 × 128 × 2 bytes = 131,072 bytes = 128 KiB per token.
  • At 8,192 tokens, one user: 8,192 × 128 KiB = 1 GiB.
  • At 32,768 tokens: 4 GiB.
  • Total at 32k context: ~4.8 + 4 + ~1.5 overhead ≈ 10–11 GB.

So a 16 GB Apple silicon laptop can run it at 8k context comfortably; 32k context is tight once the OS and apps take their share.

Worked example 2: a team server

A Karachi fintech wants to serve a 32B dense model (assume 64 layers, 8 KV heads, head_dim 128) to 20 concurrent analysts with 8k context each, on one GPU.

  • Weights, FP8: 32 GB. Weights, AWQ 4-bit: ~18 GB.
  • KV per token (FP16): 2 × 64 × 8 × 128 × 2 = 262,144 bytes = 256 KiB.
  • KV for 20 × 8,192 tokens: 163,840 tokens × 256 KiB = 40 GiB.
  • Total (FP8 weights): 32 + 40 + ~4 ≈ 76 GB → an 80 GB GPU at full load is too tight for comfort.
  • Total (4-bit weights, 8-bit KV): 18 + 20 + ~4 ≈ 42 GB → fits on a 48 GB workstation card or comfortably on 80 GB.

In practice, not every analyst uses 8k tokens at the same moment. Servers like vLLM allocate KV cache in pages on demand (PagedAttention), so real concurrency is usually higher than this worst case. Size for the worst case you care about, then load-test.

Speed: the second half of sizing

For single-user generation, a useful ceiling is:

tokens/second (decode, batch 1) ≲ memory_bandwidth ÷ bytes_read_per_token
bytes_read_per_token ≈ weights size (dense) or active-parameter weights (MoE)

Example (illustrative): a GPU with ~1,000 GB/s bandwidth running a 4.8 GB model has a theoretical ceiling around 200 tokens/s; real numbers are lower. This also explains why MoE models generate fast: they read only the active experts per token. Prompt processing ("prefill") is compute-bound instead, which is why long prompts take noticeable time before the first token.

Hands-on: a sizing calculator

# vram_calc.py: rough memory estimate for LLM inference (illustrative, verify on hardware)
import json, sys

def estimate(params_b, bits, layers, kv_heads, head_dim, ctx, users, kv_bytes=2, overhead_gb=2.0):
    weights_gb = params_b * 1e9 * bits / 8 / 1e9
    kv_per_token = 2 * layers * kv_heads * head_dim * kv_bytes
    kv_gb = kv_per_token * ctx * users / 1e9
    return weights_gb, kv_gb, weights_gb + kv_gb + overhead_gb

if __name__ == "__main__":
    # usage: python vram_calc.py config.json 8.0 4.8 8192 4
    cfg = json.load(open(sys.argv[1]))
    params_b, bits, ctx, users = float(sys.argv[2]), float(sys.argv[3]), int(sys.argv[4]), int(sys.argv[5])
    layers = cfg["num_hidden_layers"]
    kv_heads = cfg.get("num_key_value_heads", cfg["num_attention_heads"])
    head_dim = cfg.get("head_dim") or cfg["hidden_size"] // cfg["num_attention_heads"]
    w, kv, total = estimate(params_b, bits, layers, kv_heads, head_dim, ctx, users)
    print(f"weights ≈ {w:.1f} GB | KV cache ≈ {kv:.1f} GB | total ≈ {total:.1f} GB (+ verify)")

Download only the model's config.json, run the calculator for three scenarios (solo, team, peak), then confirm with the runtime's own report (for example ollama ps, llama.cpp's load log, or vLLM's startup log showing KV cache capacity).

Pitfalls

  • Using attention heads instead of KV heads (overstates KV cache for GQA models).
  • Forgetting concurrency: KV cache multiplies by simultaneous sequences.
  • Sizing to 100% of memory. Leave headroom; the OS, drivers and fragmentation need space.
  • Assuming advertised maximum context is affordable. Offer the context your users need.

How to measure success

Your estimate lands within roughly 10–20% of the runtime's reported usage, and your load test confirms the concurrency you planned for.

Key takeaways

  • Total memory ≈ weights + KV cache + overhead
  • Weights ≈ parameters × effective bits ÷ 8; use total parameters for MoE
  • KV cache ≈ 2 × layers × KV heads × head dim × bytes × tokens × concurrent sequences
  • Decode speed is roughly bounded by memory bandwidth ÷ bytes read per token
  • Estimate on paper, then confirm with the runtime report and a load test

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A model has 32 layers, 8 KV heads, head_dim 128, FP16 KV cache. What is the KV cache per token?
  2. Your estimate ignores concurrency. What will happen when 10 users hit the server?
  3. Why do MoE models often generate quickly relative to total size?

Put it into practice

Fetch a model's config.json, run the sizing calculator for solo, team and peak scenarios, then compare with your runtime's reported memory.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.