Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsServing at scale · Lesson 11 of 16
Choosing a serving engine: SGLang, TGI, TensorRT-LLM and friends
Video lecture
Choosing a serving engine: SGLang, TGI, TensorRT-LLM and friends
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Choosing a serving engine
Every few months someone posts a chart claiming a new serving engine is three times faster. Sometimes it is, for their traffic, on their hardware. The question is whether it is faster for yours. In this lesson you will get a clear map of the main engines as of September twenty twenty-six, a five-question decision guide, and a small bake-off script to test engines with your own prompts.
0:30 Analogy: courier companies
Choosing a serving engine is like choosing a courier company. They all deliver parcels. One has the biggest network and works almost everywhere. One is especially fast on repeat routes to the same addresses. One is a specialist with custom vans that are fastest on one type of road but take longer to set up. And one is a local courier that goes anywhere, even without roads. You pick based on your routes and your parcels, and you test with real deliveries before signing a contract.
1:07 The landscape (Sep 2026)
The main options. vLLM is the default for GPU serving: broad model and hardware support and a huge community. SGLang is a fast runtime whose RadixAttention reuses shared prompt prefixes automatically, and it is strong for multi-turn chat, agents and structured output. TensorRT-LLM squeezes maximum performance from NVIDIA GPUs, with more build complexity. The llama dot cpp server runs almost anywhere. Ollama is the simplest to manage. And Hugging Face's TGI, once very popular, was moved into maintenance mode in late twenty twenty-five, with Hugging Face pointing new deployments to vLLM or SGLang.
1:47 Shared prefixes
Here is what RadixAttention buys you. Imagine a support assistant where every request starts with the same three-thousand-token system prompt and often the same popular policy documents. SGLang keeps the attention cache for those shared beginnings in a tree and reuses it, so it only computes each user's unique part. vLLM has prefix caching too; the implementations differ, and so do the workloads where each shines. That is exactly why you test on your own traffic.
2:20 Five questions
Five questions decide it. One, hardware: NVIDIA data-center cards open every option, AMD needs a check of current support, Apple and CPUs point to llama dot cpp. Two, model support: does the engine handle your exact architecture, quantization, tool-call format and reasoning output? Three, traffic shape: long shared prefixes, or many small independent requests? Four, team skills: do you have GPU and Kubernetes operators, or should you start with Docker or a managed endpoint? Five, longevity: active releases and security fixes. Do not start new work on maintenance-mode software.
2:59 Managed option
Remember you do not always have to run GPUs yourself. Managed endpoints and inference clouds host open-weight models behind an API. You keep the benefits of open weights, like version choice and no lock-in to one model vendor, but your data now travels to that provider. So check the region and the data terms, especially if you operate under Saudi, Emirati, Pakistani or UK data protection rules.
3:28 Bake-off rules
Now the bake-off. Same model, same precision, same context limit, same prompts, same request rates, on both engines. The lesson includes a short async Python script that streams responses from each engine at a fixed concurrency and prints median and ninety-fifth percentile time to first token. For rigorous numbers, use each engine's own benchmark tool. For a quick apples-to-apples feel with your real prompts, this is perfect, and because both speak the OpenAI API, switching is just a URL.
4:02 Worked example: migrating off TGI
A migration story. A Manchester e-commerce firm has run TGI since twenty twenty-four. They want newer models that it will not support. So they stand vLLM up beside TGI behind their gateway, mirror five percent of traffic for two weeks, compare quality and latency, then switch. Because both expose OpenAI-style endpoints, the application change is a configuration flip. That is the pattern for any engine change: shadow, compare, switch.
4:32 Simple example: university Q&A
A simple example. A university IT team in Islamabad runs an internal Q and A bot on one GPU. Every request starts with the same long policy document. They try vLLM with prefix caching on and SGLang with its default prefix reuse, same model, same precision, same prompts. For their traffic, both give similar first-token times, and vLLM fits their existing Docker setup better, so they pick vLLM. The bake-off took one afternoon and ended a two-week debate.
5:06 Check the parsers
One more consideration people forget: how the engine handles the model's special outputs. Reasoning models may emit a thinking section before the answer. Tool-calling models format calls in family-specific ways. Engines use parsers to turn these into clean API fields, and if the parser for your model is missing or buggy, your application sees garbled tool calls or leaked reasoning text. Before committing, test your exact model's tool calls and reasoning output through each engine's API, not just plain chat.
5:41 FAQ
People often ask: will I be locked into one engine? Much less than you might fear. Because vLLM, SGLang, llama dot cpp and most managed endpoints expose OpenAI-compatible APIs, your application code rarely changes. The real switching costs are operational: container images, monitoring dashboards, flags and runbooks. Keep those documented and your engine becomes a replaceable component. Another question: should small teams even run their own engine? If you have steady volume, data residency needs or cost pressure, yes. If not, a managed endpoint may be the better first step.
6:20 Try this now
Try this now. List the three questions that would decide your engine choice today: your hardware, your exact model and features like tool calling, and whether your traffic shares long prefixes. Write one sentence per question. If any answer is I do not know, that is the first thing to find out before you compare engines.
6:44 Watch me do it
Watch me do it. I start vLLM on port eight thousand and SGLang on port thirty thousand, same model, same precision, same maximum context. I check both health endpoints. Then I prepare one hundred real prompts, most of which share our long system prompt, and run the bake-off script at sixteen concurrent requests against each engine. It prints median and ninety-fifth percentile time to first token. I repeat at thirty-two concurrent requests. Next I test what the lesson warned about: tool calls. I send ten prompts that should trigger our lookup tool and check both engines return properly parsed tool calls. One engine needs a specific parser flag for this model family; I add it and rerun. Finally I write the table: engine, first-token latency at both loads, error count, tool-call success, and a sentence on which one fits our Docker setup. The decision takes an afternoon and ends the debate.
7:50 Recap
Recap. vLLM is the sensible default for GPU serving, SGLang shines with shared prefixes and structured output, TensorRT-LLM for peak NVIDIA performance, llama dot cpp for anywhere else, and TGI only for existing deployments you plan to migrate. Decide on hardware, model support, traffic, skills and longevity, then prove it with a fair bake-off. Your next step: run the bake-off on your hardware with a hundred of your prompts.
There is no single best engine
Serving engines differ in model coverage, hardware support, features and operational maturity. The right choice depends on your hardware, traffic shape, model family and team. This lesson maps the landscape as of September 2026 and gives you a way to decide. Versions move quickly, so confirm current support for your exact model in each project's documentation.
The main options
| Engine | Strengths | Typical fit |
|---|---|---|
| vLLM | Broad model and hardware support, large community, PagedAttention, continuous batching, OpenAI-compatible, multi-LoRA | Default choice for GPU serving |
| SGLang | Fast serving runtime with RadixAttention (automatic reuse of shared prefixes across requests via a radix tree), efficient structured outputs, strong multi-turn and agentic performance | Workloads with heavy shared prefixes (RAG, multi-turn chat, agents), structured output at scale |
| TensorRT-LLM (NVIDIA) | Highly optimized kernels for NVIDIA GPUs; used under NVIDIA's inference products | Squeezing maximum performance on NVIDIA hardware, with more build complexity |
| llama.cpp server | Runs almost anywhere (CPU, Apple, consumer GPUs), GGUF, tiny footprint | Edge servers, small teams, mixed hardware |
| Ollama | Simplest management and developer experience | Individuals and small teams, prototypes |
| Hugging Face TGI | Historically popular; Hugging Face moved TGI into maintenance mode (announced late 2025) and points users to vLLM or SGLang for new deployments | Existing deployments only; plan migration |
Managed options also exist: Hugging Face Inference Endpoints, cloud providers' model catalogs, and specialized inference clouds that host open-weight models behind an API. These give you open weights without running GPUs yourself, though your data then goes to that provider (check their region and data terms).
How SGLang differs, in one example
A RAG assistant sends the same 3,000-token system prompt plus the same popular policy documents to thousands of users. RadixAttention keeps the KV cache for those shared prefixes in a tree and reuses it automatically, so only each user's unique suffix needs fresh computation. vLLM also offers prefix caching; the engines differ in implementation and in which workloads benefit most. Benchmark both on your traffic.
# SGLang: launch an OpenAI-compatible server (see SGLang docs for install and flags)
python -m sglang.launch_server --model-path Qwen/Qwen3-8B --port 30000 --api-key "$SGLANG_API_KEY"Your client code is unchanged apart from the base URL, which makes A/B testing engines cheap.
A decision guide
- Hardware first. NVIDIA data-center GPUs: vLLM, SGLang or TensorRT-LLM. AMD accelerators: check vLLM and SGLang ROCm support for your model. Apple or CPU-heavy: llama.cpp (or Ollama/LM Studio).
- Model support. Does the engine support your exact architecture, quantization and features (tool-call parser, vision input, reasoning output parsing)? New architectures often land in vLLM and SGLang first.
- Traffic shape. Shared long prefixes and multi-turn: test SGLang and vLLM prefix caching. Many small independent requests: either works; measure.
- Team skills. Kubernetes and GPU operations? Any engine. Small IT team? Start with vLLM Docker or a managed endpoint.
- Support and longevity. Active releases, security fixes, community. Avoid starting new work on engines in maintenance mode.
Hands-on: an engine bake-off
Run two engines with the same model, same precision and same context limit, then replay the same prompt set at the same request rates.
# bakeoff.py: minimal concurrent load generator (pip install openai)
import asyncio, os, time, statistics
from openai import AsyncOpenAI
ENGINES = {"vllm": "http://localhost:8000/v1", "sglang": "http://localhost:30000/v1"}
PROMPTS = [line.strip() for line in open("prompts.txt", encoding="utf-8") if line.strip()]
async def one(client, model, prompt):
start = time.perf_counter(); first = None
stream = await client.chat.completions.create(model=model, stream=True, max_tokens=256,
messages=[{"role": "user", "content": prompt}])
async for chunk in stream:
if first is None and chunk.choices and chunk.choices[0].delta.content:
first = time.perf_counter() - start
return first, time.perf_counter() - start
async def run(name, url, model, concurrency=16):
client = AsyncOpenAI(base_url=url, api_key=os.getenv("ENGINE_API_KEY", "none"), timeout=120)
sem = asyncio.Semaphore(concurrency)
async def guarded(p):
async with sem:
try:
return await one(client, model, p)
except Exception as e:
print(name, "error:", e); return None
results = [r for r in await asyncio.gather(*(guarded(p) for p in PROMPTS)) if r and r[0]]
ttft = sorted(r[0] for r in results)
print(f"{name}: n={len(results)} p50 TTFT={statistics.median(ttft):.2f}s p95 TTFT={ttft[int(0.95*len(ttft))-1]:.2f}s")
async def main():
for name, url in ENGINES.items():
await run(name, url, os.getenv("MODEL", "Qwen/Qwen3-8B"))
asyncio.run(main())Use the engines' own benchmark tools for rigorous numbers; this script is for a quick, apples-to-apples feel with your prompts.
Worked example: migrating off TGI
A Manchester e-commerce firm has run TGI since 2024. New models they want are not supported, and TGI is in maintenance mode. They stand up vLLM beside TGI behind their gateway, mirror 5% of traffic, compare outputs and latency for two weeks, then switch. Because both expose OpenAI-style endpoints, application changes are a config flip.
Pitfalls
- Comparing engines at different precisions or context limits.
- Choosing on a vendor benchmark with a traffic shape unlike yours.
- Adopting a new engine without checking tool-call and reasoning parsers for your model.
- Starting new projects on maintenance-mode software.
How to measure success
A bake-off table (same model, precision, context, prompts, rates) with p50/p95 TTFT, throughput and error rate, plus a short rationale covering hardware, model support and team skills.
Key takeaways
- vLLM is the default GPU choice; SGLang excels at shared prefixes and structured output
- TensorRT-LLM maximizes NVIDIA performance; llama.cpp runs almost anywhere
- Hugging Face TGI is in maintenance mode; plan migrations and avoid it for new work
- Decide on hardware, model support, traffic shape, team skills and longevity
- Bake off engines on the same model, precision, context and prompts
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run a bake-off between two engines on your hardware with 100 of your prompts at two concurrency levels; record p50/p95 TTFT, throughput and errors.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.