Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsServing at scale · Lesson 10 of 16
Serving at scale with vLLM
Video lecture
Serving at scale with vLLM
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Serving at scale with vLLM
Your pilot was a hit. Twelve people loved the local assistant. Now the whole company wants it, plus an application that fires hundreds of requests at once. The desktop runner that was perfect for one person starts to queue, stall and time out. This lesson is about serving at scale with vLLM: why it is fast, how to launch it properly, which flags to tune, and how to measure performance like an operator.
0:32 Analogy: a professional kitchen
Think of a busy restaurant again. A desktop runner is a single chef cooking one order at a time from start to finish. vLLM is a professional kitchen with a head chef who starts new dishes the moment a burner frees up, shares prep work between similar orders, and never leaves a stove idle. That is continuous batching and prefix caching. And PagedAttention is smart shelving: instead of reserving a whole shelf for each order, ingredients go wherever there is space. Same ingredients, same chefs, far more dishes per hour.
1:11 Why vLLM is fast
Two ideas make vLLM fast. The first is PagedAttention. Instead of reserving one huge block of memory per request for context it might never use, vLLM stores the attention cache in small pages allocated on demand, like virtual memory in an operating system. Less waste, more requests per GPU. The second is continuous batching. Requests join and leave the running batch at every step, so a short answer never waits for a long one, and the GPU never sits idle.
1:46 Production features
On top of that you get prefix caching, which reuses work for shared prompt beginnings like a long system prompt or the same retrieved document. You get quantized checkpoints such as FP8, AWQ and GPTQ, tensor parallelism across GPUs, structured outputs, tool calling, and serving many fine-tuned adapters on one base model. And it speaks the OpenAI API, so your existing code barely changes.
2:14 Launch and tune
Launching is one command: vllm serve, the model name, a maximum context length, a GPU memory budget and an API key. Or use the official Docker image, pinned to a specific tag. Then watch the startup log. vLLM tells you how much memory it allocated for the cache and how many requests it can hold at full context. That single line is your capacity plan. Two flags matter most. Max model length: set it to what users need, not the model's maximum, because every extra token of allowed context reduces concurrency. And GPU memory utilization: lower it if something else shares the card.
2:59 Measure it
Now measure like an operator. Users feel two things: time to first token, how long before text appears, and inter-token latency, how smoothly it streams. Operators also watch throughput and the ninety-fifth percentile, because averages hide the unhappy users. vLLM includes a benchmarking command that replays prompts at a target rate and reports all of these. And it exposes Prometheus metrics, queue length, cache usage and token throughput, ready for dashboards and alerts.
3:31 Worked example: Abu Dhabi insurer (illustrative)
Let's walk through an insurer in Abu Dhabi, with illustrative numbers. A hundred and twenty claims adjusters, around forty at peak. They serve a thirty-two billion model at FP8 on one eighty gigabyte GPU with a twelve thousand token context. They replay five hundred real, redacted prompts at two, four and eight requests per second. At four per second, the ninety-fifth percentile first token is under two seconds. Fine. At eight, it passes six seconds. Not fine. So one GPU handles normal peaks, and a second replica handles month-end surges. Alerts fire when the queue or cache usage runs hot.
4:14 Production checklist
Before you call it production, run the checklist. Authentication and TLS, and network isolation. Pinned model revision and container tag, recorded in your register. Health checks. Metrics, dashboards and alerts. A load test before launch and after every model change. Rate limits per client at the gateway. And logging that respects your privacy policy. Three classic mistakes: maximum context set to the model limit, a so-called load test with one request at a time, and an unauthenticated server because it is only internal.
4:50 Simple example: 10-person firm
A simple example. A ten-person accountancy firm in Birmingham runs vLLM on one workstation GPU with an eight billion model. They set maximum context to eight thousand tokens because their documents are short after retrieval. At startup, the log says the cache can hold dozens of eight-thousand-token requests at once, far more than ten staff will ever send together. They run a quick benchmark at two requests per second, see first tokens in well under a second, and they are done. Not every deployment needs a capacity war room.
5:29 Scale up vs scale out
A word on scaling out. When one GPU is not enough, you have two options. Scale up: a bigger card, or tensor parallelism across several cards in one machine, for models that do not fit. Or scale out: several identical replicas behind a load balancer, for more concurrent users. For most business workloads, replicas are simpler, more resilient, and let you do rolling upgrades, one replica at a time, with no downtime. Decide your scale-out trigger in advance, for example, queue length above a threshold for ten minutes.
6:07 FAQ: users per GPU?
A frequent question: how many users can one GPU serve? It depends on four things: model size and precision, the context length you allow, how long answers are, and how often people actually send requests. Ten people who each send a request every few minutes generate far less load than ten people hammering a chat window. The honest answer is to measure: replay realistic traffic at increasing request rates and find the point where your latency target breaks. That number, not a guess, goes into your capacity plan.
6:45 Try this now
Try this now. If you have access to any NVIDIA GPU, even in a notebook environment, launch vLLM with a small model and a maximum context of four thousand tokens. Find the log line that reports the KV cache capacity. Then restart with sixteen thousand and read it again. Watching that number shrink is the fastest way to understand why context limits matter for concurrency.
7:13 Watch me do it
Watch me do it. I launch vllm serve with an eight billion model, maximum context eight thousand, memory utilization zero point nine and an API key. While it starts, I watch the log and find the line reporting KV cache capacity and maximum concurrency. I note both numbers. Then I run the benchmark command, replaying two hundred prompts from our support logs at two, four and eight requests per second. For each rate I copy the median and ninety-fifth percentile time to first token and inter-token latency into a table. At four requests per second, first tokens arrive in about a second at p ninety-five. At eight, queueing pushes it past four seconds. I open the metrics endpoint and point a dashboard at it, with an alert when queue length stays high for five minutes. Then I write the capacity note: one GPU handles up to four requests per second within target; add a replica beyond that. Illustrative numbers, real method.
8:23 Recap
Recap. vLLM is built for throughput under concurrency thanks to PagedAttention and continuous batching. Tune maximum context and memory first, and read the capacity line in the log. Measure first-token and inter-token latency at realistic load, watch the ninety-fifth percentile, and monitor metrics. Your next step: serve a model, benchmark it at three request rates with your own prompts, and write a one-page capacity note.
Why a desktop runner is not a server
Ollama and LM Studio are brilliant for one person. When 50 people or an application with bursts of traffic hit the same model, you need an engine designed for throughput under concurrency. vLLM is the most widely used open-source serving engine for this. It exposes an OpenAI-compatible API, supports most popular open-weight architectures, and runs on NVIDIA and AMD GPUs among other hardware (check its docs for your accelerator).
The two ideas that make vLLM fast
- PagedAttention. Instead of reserving one big contiguous block of KV cache per request (wasting memory on context that may never be used), vLLM stores KV cache in small pages allocated on demand, like virtual memory in an operating system. Less waste means more concurrent requests per GPU.
- Continuous batching. Requests join and leave the running batch at every generation step rather than waiting for a whole batch to finish. Short answers do not wait for long ones; the GPU stays busy.
Add prefix caching (reusing KV cache for shared prompt prefixes such as a long system prompt or the same retrieved document), quantized checkpoints (FP8, AWQ, GPTQ and others), tensor parallelism across GPUs, structured outputs, tool calling and multi-LoRA serving, and you have a production engine.
Hands-on: serve a model
# Python install (see docs for CUDA/ROCm specifics)
pip install vllm
# Serve with an API key, a context limit and a memory budget
export VLLM_API_KEY="$(openssl rand -hex 24)"
vllm serve Qwen/Qwen3-8B \
--max-model-len 16384 \
--gpu-memory-utilization 0.90 \
--api-key "$VLLM_API_KEY" \
--port 8000Or with the official Docker image:
docker run --gpus all -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-e HF_TOKEN="$HF_TOKEN" \
vllm/vllm-openai:latest \
--model Qwen/Qwen3-8B --max-model-len 16384 --api-key "$VLLM_API_KEY"Pin a specific image tag in production rather than latest. Then call it like any OpenAI-compatible server:
import os
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key=os.environ["VLLM_API_KEY"])
r = client.chat.completions.create(model="Qwen/Qwen3-8B",
messages=[{"role": "user", "content": "Summarize PagedAttention in one line."}])
print(r.choices[0].message.content)The flags you will actually tune
| Flag | What it controls | Rule of thumb |
|---|---|---|
--max-model-len | Maximum context per request | Set to what users need, not the model maximum; more context means fewer concurrent requests |
--gpu-memory-utilization | Fraction of GPU memory vLLM may use (default 0.9) | Lower it if other processes share the GPU |
--tensor-parallel-size | Split the model across N GPUs | For models that do not fit on one GPU |
--kv-cache-dtype | KV cache precision (e.g. fp8) | Roughly doubles KV capacity vs 16-bit, small quality risk; test |
--enable-lora, --lora-modules | Serve fine-tuned adapters on one base model | See the fine-tuning course |
--enable-auto-tool-choice, --tool-call-parser | Tool calling for models that support it | Parser must match the model family |
At startup vLLM logs how much KV cache it allocated and the implied maximum concurrency at full context. Read that line: it is your capacity plan in one number.
Measuring performance properly
Users feel two latencies:
- TTFT (time to first token): how long until text starts streaming. Dominated by queueing and prompt processing.
- ITL/TPOT (inter-token latency / time per output token): how fast text streams once it starts.
Operators care about throughput (requests or tokens per second) and p95/p99 latencies under realistic concurrency. vLLM ships a benchmarking CLI (vllm bench serve in recent versions; check vllm bench --help) that replays prompts at a target request rate and reports these metrics. vLLM also exposes Prometheus metrics at /metrics (queue length, KV cache usage, token throughput) for Grafana dashboards and alerts.
Worked example: a UAE insurer's claims assistant
An insurer in Abu Dhabi serves an internal claims assistant to 120 adjusters, peak ~40 concurrent. Steps (numbers illustrative):
- Size: a 32B model at FP8 on one 80 GB GPU;
--max-model-len 12288because claims packets rarely exceed that after retrieval. - Benchmark: replay 500 real (redacted) prompts at 2, 4 and 8 requests/second. At 4 req/s, p95 TTFT is 1.8 s and p95 ITL 45 ms: acceptable. At 8 req/s, queueing pushes p95 TTFT past 6 s: not acceptable.
- Decide: one GPU covers normal peaks; a second replica behind a load balancer covers month-end surges.
- Operate: alerts on queue length and KV cache usage above 90%; weekly review of the metrics dashboard.
Production checklist
- API key or gateway authentication, TLS, network isolation.
- Pinned model revision and container tag; model register entry.
- Health checks and readiness probes (the server exposes a health endpoint).
- Prometheus metrics, dashboards and alerts.
- Load test before launch and after every model change.
- Rate limits per client at the gateway.
- Log prompts only where policy allows; redact PII.
Pitfalls
- Setting
--max-model-lento the model maximum and wondering why concurrency collapsed. - Benchmarking with one request at a time and calling it "load testing".
- Leaving the server unauthenticated on an internal network ("it is internal" is not a control).
How to measure success
You have a load-test report showing p95 TTFT and ITL at your expected and peak request rates, a dashboard with alerts, and a documented scale-out trigger.
Key takeaways
- vLLM is built for throughput under concurrency: PagedAttention plus continuous batching
- Prefix caching, quantized checkpoints, tensor parallelism and multi-LoRA are production features
- Tune --max-model-len and --gpu-memory-utilization first; read the KV cache capacity log line
- Measure TTFT, inter-token latency, throughput and p95 under realistic load
- Authenticate, pin versions, monitor /metrics and load-test after every change
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Serve a model with vLLM, run a benchmark at three request rates using your own prompts, and write a one-page capacity note with p95 TTFT and ITL.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.