Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsServing at scale · Lesson 10 of 16

Serving at scale with vLLM

Article · 16 min · 9 min lecture

Video lecture

Serving at scale with vLLM

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Serving at scale with vLLM

  • From one user to hundreds
  • Why it is fast
  • Flags, metrics, capacity

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why a desktop runner is not a server

Ollama and LM Studio are brilliant for one person. When 50 people or an application with bursts of traffic hit the same model, you need an engine designed for throughput under concurrency. vLLM is the most widely used open-source serving engine for this. It exposes an OpenAI-compatible API, supports most popular open-weight architectures, and runs on NVIDIA and AMD GPUs among other hardware (check its docs for your accelerator).

The two ideas that make vLLM fast

  1. PagedAttention. Instead of reserving one big contiguous block of KV cache per request (wasting memory on context that may never be used), vLLM stores KV cache in small pages allocated on demand, like virtual memory in an operating system. Less waste means more concurrent requests per GPU.
  2. Continuous batching. Requests join and leave the running batch at every generation step rather than waiting for a whole batch to finish. Short answers do not wait for long ones; the GPU stays busy.

Add prefix caching (reusing KV cache for shared prompt prefixes such as a long system prompt or the same retrieved document), quantized checkpoints (FP8, AWQ, GPTQ and others), tensor parallelism across GPUs, structured outputs, tool calling and multi-LoRA serving, and you have a production engine.

Hands-on: serve a model

# Python install (see docs for CUDA/ROCm specifics)
pip install vllm

# Serve with an API key, a context limit and a memory budget
export VLLM_API_KEY="$(openssl rand -hex 24)"
vllm serve Qwen/Qwen3-8B \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.90 \
  --api-key "$VLLM_API_KEY" \
  --port 8000

Or with the official Docker image:

docker run --gpus all -p 8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -e HF_TOKEN="$HF_TOKEN" \
  vllm/vllm-openai:latest \
  --model Qwen/Qwen3-8B --max-model-len 16384 --api-key "$VLLM_API_KEY"

Pin a specific image tag in production rather than latest. Then call it like any OpenAI-compatible server:

import os
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key=os.environ["VLLM_API_KEY"])
r = client.chat.completions.create(model="Qwen/Qwen3-8B",
                                   messages=[{"role": "user", "content": "Summarize PagedAttention in one line."}])
print(r.choices[0].message.content)

The flags you will actually tune

FlagWhat it controlsRule of thumb
--max-model-lenMaximum context per requestSet to what users need, not the model maximum; more context means fewer concurrent requests
--gpu-memory-utilizationFraction of GPU memory vLLM may use (default 0.9)Lower it if other processes share the GPU
--tensor-parallel-sizeSplit the model across N GPUsFor models that do not fit on one GPU
--kv-cache-dtypeKV cache precision (e.g. fp8)Roughly doubles KV capacity vs 16-bit, small quality risk; test
--enable-lora, --lora-modulesServe fine-tuned adapters on one base modelSee the fine-tuning course
--enable-auto-tool-choice, --tool-call-parserTool calling for models that support itParser must match the model family

At startup vLLM logs how much KV cache it allocated and the implied maximum concurrency at full context. Read that line: it is your capacity plan in one number.

Measuring performance properly

Users feel two latencies:

  • TTFT (time to first token): how long until text starts streaming. Dominated by queueing and prompt processing.
  • ITL/TPOT (inter-token latency / time per output token): how fast text streams once it starts.

Operators care about throughput (requests or tokens per second) and p95/p99 latencies under realistic concurrency. vLLM ships a benchmarking CLI (vllm bench serve in recent versions; check vllm bench --help) that replays prompts at a target request rate and reports these metrics. vLLM also exposes Prometheus metrics at /metrics (queue length, KV cache usage, token throughput) for Grafana dashboards and alerts.

Worked example: a UAE insurer's claims assistant

An insurer in Abu Dhabi serves an internal claims assistant to 120 adjusters, peak ~40 concurrent. Steps (numbers illustrative):

  1. Size: a 32B model at FP8 on one 80 GB GPU; --max-model-len 12288 because claims packets rarely exceed that after retrieval.
  2. Benchmark: replay 500 real (redacted) prompts at 2, 4 and 8 requests/second. At 4 req/s, p95 TTFT is 1.8 s and p95 ITL 45 ms: acceptable. At 8 req/s, queueing pushes p95 TTFT past 6 s: not acceptable.
  3. Decide: one GPU covers normal peaks; a second replica behind a load balancer covers month-end surges.
  4. Operate: alerts on queue length and KV cache usage above 90%; weekly review of the metrics dashboard.

Production checklist

  • API key or gateway authentication, TLS, network isolation.
  • Pinned model revision and container tag; model register entry.
  • Health checks and readiness probes (the server exposes a health endpoint).
  • Prometheus metrics, dashboards and alerts.
  • Load test before launch and after every model change.
  • Rate limits per client at the gateway.
  • Log prompts only where policy allows; redact PII.

Pitfalls

  • Setting --max-model-len to the model maximum and wondering why concurrency collapsed.
  • Benchmarking with one request at a time and calling it "load testing".
  • Leaving the server unauthenticated on an internal network ("it is internal" is not a control).

How to measure success

You have a load-test report showing p95 TTFT and ITL at your expected and peak request rates, a dashboard with alerts, and a documented scale-out trigger.

Key takeaways

  • vLLM is built for throughput under concurrency: PagedAttention plus continuous batching
  • Prefix caching, quantized checkpoints, tensor parallelism and multi-LoRA are production features
  • Tune --max-model-len and --gpu-memory-utilization first; read the KV cache capacity log line
  • Measure TTFT, inter-token latency, throughput and p95 under realistic load
  • Authenticate, pin versions, monitor /metrics and load-test after every change

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What does PagedAttention primarily improve?
  2. Users complain that answers take long to start but then stream quickly. Which metric is high?
  3. Concurrency collapsed after you set --max-model-len to the model maximum. Why?

Put it into practice

Serve a model with vLLM, run a benchmark at three request rates using your own prompts, and write a one-page capacity note with p95 TTFT and ITL.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.