Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsServing at scale · Lesson 11 of 16

Choosing a serving engine: SGLang, TGI, TensorRT-LLM and friends

Article · 14 min · 8 min lecture

Video lecture

Choosing a serving engine: SGLang, TGI, TensorRT-LLM and friends

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Choosing a serving engine

  • The landscape
  • A five-question guide
  • Your own bake-off

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

There is no single best engine

Serving engines differ in model coverage, hardware support, features and operational maturity. The right choice depends on your hardware, traffic shape, model family and team. This lesson maps the landscape as of September 2026 and gives you a way to decide. Versions move quickly, so confirm current support for your exact model in each project's documentation.

The main options

EngineStrengthsTypical fit
vLLMBroad model and hardware support, large community, PagedAttention, continuous batching, OpenAI-compatible, multi-LoRADefault choice for GPU serving
SGLangFast serving runtime with RadixAttention (automatic reuse of shared prefixes across requests via a radix tree), efficient structured outputs, strong multi-turn and agentic performanceWorkloads with heavy shared prefixes (RAG, multi-turn chat, agents), structured output at scale
TensorRT-LLM (NVIDIA)Highly optimized kernels for NVIDIA GPUs; used under NVIDIA's inference productsSqueezing maximum performance on NVIDIA hardware, with more build complexity
llama.cpp serverRuns almost anywhere (CPU, Apple, consumer GPUs), GGUF, tiny footprintEdge servers, small teams, mixed hardware
OllamaSimplest management and developer experienceIndividuals and small teams, prototypes
Hugging Face TGIHistorically popular; Hugging Face moved TGI into maintenance mode (announced late 2025) and points users to vLLM or SGLang for new deploymentsExisting deployments only; plan migration

Managed options also exist: Hugging Face Inference Endpoints, cloud providers' model catalogs, and specialized inference clouds that host open-weight models behind an API. These give you open weights without running GPUs yourself, though your data then goes to that provider (check their region and data terms).

How SGLang differs, in one example

A RAG assistant sends the same 3,000-token system prompt plus the same popular policy documents to thousands of users. RadixAttention keeps the KV cache for those shared prefixes in a tree and reuses it automatically, so only each user's unique suffix needs fresh computation. vLLM also offers prefix caching; the engines differ in implementation and in which workloads benefit most. Benchmark both on your traffic.

# SGLang: launch an OpenAI-compatible server (see SGLang docs for install and flags)
python -m sglang.launch_server --model-path Qwen/Qwen3-8B --port 30000 --api-key "$SGLANG_API_KEY"

Your client code is unchanged apart from the base URL, which makes A/B testing engines cheap.

A decision guide

  1. Hardware first. NVIDIA data-center GPUs: vLLM, SGLang or TensorRT-LLM. AMD accelerators: check vLLM and SGLang ROCm support for your model. Apple or CPU-heavy: llama.cpp (or Ollama/LM Studio).
  2. Model support. Does the engine support your exact architecture, quantization and features (tool-call parser, vision input, reasoning output parsing)? New architectures often land in vLLM and SGLang first.
  3. Traffic shape. Shared long prefixes and multi-turn: test SGLang and vLLM prefix caching. Many small independent requests: either works; measure.
  4. Team skills. Kubernetes and GPU operations? Any engine. Small IT team? Start with vLLM Docker or a managed endpoint.
  5. Support and longevity. Active releases, security fixes, community. Avoid starting new work on engines in maintenance mode.

Hands-on: an engine bake-off

Run two engines with the same model, same precision and same context limit, then replay the same prompt set at the same request rates.

# bakeoff.py: minimal concurrent load generator (pip install openai)
import asyncio, os, time, statistics
from openai import AsyncOpenAI

ENGINES = {"vllm": "http://localhost:8000/v1", "sglang": "http://localhost:30000/v1"}
PROMPTS = [line.strip() for line in open("prompts.txt", encoding="utf-8") if line.strip()]

async def one(client, model, prompt):
    start = time.perf_counter(); first = None
    stream = await client.chat.completions.create(model=model, stream=True, max_tokens=256,
                                                  messages=[{"role": "user", "content": prompt}])
    async for chunk in stream:
        if first is None and chunk.choices and chunk.choices[0].delta.content:
            first = time.perf_counter() - start
    return first, time.perf_counter() - start

async def run(name, url, model, concurrency=16):
    client = AsyncOpenAI(base_url=url, api_key=os.getenv("ENGINE_API_KEY", "none"), timeout=120)
    sem = asyncio.Semaphore(concurrency)
    async def guarded(p):
        async with sem:
            try:
                return await one(client, model, p)
            except Exception as e:
                print(name, "error:", e); return None
    results = [r for r in await asyncio.gather(*(guarded(p) for p in PROMPTS)) if r and r[0]]
    ttft = sorted(r[0] for r in results)
    print(f"{name}: n={len(results)} p50 TTFT={statistics.median(ttft):.2f}s p95 TTFT={ttft[int(0.95*len(ttft))-1]:.2f}s")

async def main():
    for name, url in ENGINES.items():
        await run(name, url, os.getenv("MODEL", "Qwen/Qwen3-8B"))

asyncio.run(main())

Use the engines' own benchmark tools for rigorous numbers; this script is for a quick, apples-to-apples feel with your prompts.

Worked example: migrating off TGI

A Manchester e-commerce firm has run TGI since 2024. New models they want are not supported, and TGI is in maintenance mode. They stand up vLLM beside TGI behind their gateway, mirror 5% of traffic, compare outputs and latency for two weeks, then switch. Because both expose OpenAI-style endpoints, application changes are a config flip.

Pitfalls

  • Comparing engines at different precisions or context limits.
  • Choosing on a vendor benchmark with a traffic shape unlike yours.
  • Adopting a new engine without checking tool-call and reasoning parsers for your model.
  • Starting new projects on maintenance-mode software.

How to measure success

A bake-off table (same model, precision, context, prompts, rates) with p50/p95 TTFT, throughput and error rate, plus a short rationale covering hardware, model support and team skills.

Key takeaways

  • vLLM is the default GPU choice; SGLang excels at shared prefixes and structured output
  • TensorRT-LLM maximizes NVIDIA performance; llama.cpp runs almost anywhere
  • Hugging Face TGI is in maintenance mode; plan migrations and avoid it for new work
  • Decide on hardware, model support, traffic shape, team skills and longevity
  • Bake off engines on the same model, precision, context and prompts

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your RAG assistant sends the same long system prompt and popular documents to thousands of users. Which feature matters most?
  2. A team plans a new deployment on Hugging Face TGI in 2026. What is the key concern?
  3. Which comparison is fair?

Put it into practice

Run a bake-off between two engines on your hardware with 100 of your prompts at two concurrency levels; record p50/p95 TTFT, throughput and errors.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.