Open-Weight and Local AI: Run, Choose and Deploy Your Own Models · Running models locally · lesson 8 of 16 · 15 min
llama.cpp and LM Studio: control and convenience
Two tools, two jobs
llama.cpp is the open-source C/C++ inference engine behind much of the local-LLM world (Ollama and LM Studio use it among their back ends). Using it directly gives you maximum control: exact quantization, GPU offload, context, sampling, grammar-constrained output and a lightweight OpenAI-compatible server.
LM Studio is a desktop app (macOS, Windows, Linux) with a model browser, chat UI and a local server. On Apple silicon it can run both GGUF (via llama.cpp) and MLX models. It is ideal for non-engineers and for quickly comparing models side by side. Check its license terms for workplace use on its website.
llama.cpp: the server in practice
Build or install llama.cpp per its README (prebuilt binaries and package managers are available). Then:
# Download a GGUF from Hugging Face and serve it (repo:quant syntax; check the model page for available quants)
llama-server -hf ggml-org/gpt-oss-20b-GGUF -c 16384 --port 8080 --jinja
# Or serve a local file with full GPU offload, API key and parallel slots
llama-server -m ./models/my-model-Q4_K_M.gguf \
-c 32768 -ngl 99 --parallel 4 \
--host 127.0.0.1 --port 8080 \
--api-key "$LLAMA_API_KEY"
Key flags:
| Flag | Meaning | |---|---| | -m / -hf | local model file / download from a Hugging Face repo | | -c | total context size (shared across parallel slots) | | -ngl | number of layers to offload to GPU (99 means "all that exist") | | --parallel | number of simultaneous sequences (slots) | | --jinja | use the model's own chat template (important for tool calling) | | --api-key | require a bearer key on requests |
Note that with --parallel 4 and -c 32768, each slot gets roughly 8k tokens. Check /health to confirm the server is ready, and use the OpenAI-compatible /v1/chat/completions from any client.
Grammar-constrained output
llama.cpp can force outputs to match a grammar (GBNF) or a JSON schema. Constrained decoding means the model literally cannot produce tokens that break the format. Useful for extraction and classification:
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Authorization: Bearer $LLAMA_API_KEY" -H "Content-Type: application/json" -d '{
"messages": [{"role":"user","content":"Classify: \"Order 5531 arrived broken, I want my money back\""}],
"response_format": {"type": "json_schema", "json_schema": {"schema": {
"type":"object",
"properties":{"intent":{"type":"string","enum":["refund","delivery","product_question","other"]},
"order_id":{"type":"string"}},
"required":["intent"]}}}
}'
Parameter names for schemas vary slightly across versions; check the server README.
Performance knobs that matter
- GPU offload (
-ngl): all layers on GPU is fastest; partial offload lets you run models larger than VRAM at reduced speed. - Flash attention and KV cache type (for example 8-bit K/V) reduce memory at long context; flags are listed in
llama-server --help. - Threads for CPU inference: usually the number of performance cores.
- Batch/ubatch sizes affect prompt-processing speed.
- Speculative decoding (a small draft model proposing tokens a big model verifies) can raise speed for some workloads.
Always benchmark with llama-bench on your hardware rather than copying someone else's flags.
LM Studio in practice
- Search and download a model in the Discover tab; the app shows which quantizations fit your memory.
- Chat with it and compare with another model in a split view.
- Start the local server (Developer tab or
lms server start); it exposes OpenAI-compatible endpoints onhttp://localhost:1234/v1by default. - Use the
lmsCLI for scripting:lms ls,lms load <model>,lms server start.
It is excellent for a marketing team that wants to try local models on real briefs before IT invests in a server.
Worked example: choosing between them
A UK charity handles safeguarding notes that must stay on-device.
- Case workers (non-technical) use LM Studio on their laptops with a small approved model, following an IT-provided setup guide and a list of approved models.
- The data team runs
llama-serveron an internal Linux box with a JSON schema that extracts risk categories from notes into their case system, with an API key and firewall rules. - IT documents both in the model register and pins model files by checksum.
Hands-on: benchmark two quantizations
llama-bench -m model-Q4_K_M.gguf -m model-Q8_0.gguf -ngl 99 -p 512 -n 128
# Reports prompt-processing (pp) and token-generation (tg) speed for each model file
Record pp and tg tokens/second next to your task scores from the quantization lesson.
Pitfalls
- Forgetting
--jinja(or the right template) and getting broken tool calls. - Setting
-cwithout accounting for--parallelslot division. - Running LM Studio's server on a shared network without understanding who can reach it.
- Downloading GGUFs from unverified uploaders.
How to measure success
You can serve a GGUF with llama.cpp behind an API key, enforce a JSON schema on output, and show llama-bench numbers that justify your quantization and offload settings.
Video lecture: llama.cpp and LM Studio: control and convenience
Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.
- llama.cpp and LM Studio
- Analogy: manual car + showroom
- llama-server essentials
- Watch out
- Constrained output
- Performance knobs
- LM Studio
- Worked example: UK charity
- Simple example: bookshop email tags
- Which one when?
- FAQ
- Try this now
- Watch me do it
- Recap
Lecture transcript
llama.cpp and LM Studio
Ollama is the easy button. But sometimes you need to open the hood: force a model to return perfect JSON, squeeze a model onto a machine that is slightly too small, or hand a friendly app to colleagues who never touch a terminal. That is where llama dot cpp and LM Studio come in. In this lesson you will run the llama dot cpp server with sensible flags, constrain output with a schema, tune performance, and set up LM Studio for non-technical teams.
Analogy: manual car + showroom
If Ollama is an automatic car, llama dot cpp is a manual with a full dashboard. You control every gear: how many layers go on the GPU, how the context is split, which template is used, even which tokens the model is allowed to choose. LM Studio is the showroom. You walk in, try models side by side, see which ones fit your garage, and drive one home without reading the manual. Different people in the same company need different vehicles.
llama-server essentials
llama dot cpp is the open-source engine under a lot of local AI tools. Its server takes a model, either a local GGUF file or one fetched directly from a Hugging Face repository, and exposes an OpenAI-compatible API. The flags you will use most: minus c for context size, minus n g l for how many layers go to the GPU, parallel for how many requests run at once, jinja to use the model's own chat template, and api key to require a key on every request.
Watch out
One subtle point. The context you set is shared across parallel slots. Thirty-two thousand tokens of context with four slots gives each request about eight thousand. If you promised users thirty-two thousand each, you will need thirty-two thousand times four, and the memory to match. And always use the jinja option, or the right template, when you rely on tool calling. A wrong template is the number one reason local tool calls look broken.
Constrained output
Now a genuinely powerful feature: constrained output. You can give the server a JSON schema, or a grammar, and the decoder will only choose tokens that keep the output valid. The model literally cannot produce broken JSON. For a support-intent classifier, you define an intent field that must be one of refund, delivery, product question or other, plus an optional order number. The format is guaranteed. Whether the label is correct still needs evaluation.
Performance knobs
Performance knobs. Put all layers on the GPU if they fit. If not, partial offload lets a too-big model run, just slower. Flash attention and an eight-bit KV cache save memory at long context. For CPU inference, match threads to performance cores. Speculative decoding, where a small draft model proposes tokens and the big model checks them, can speed some workloads. But do not copy flags from forums. Run llama bench on your own machine and compare.
LM Studio
LM Studio is the friendly face. You search for a model, and it tells you which versions fit your memory. You chat, compare two models side by side, and when you are ready, start a local server that speaks the OpenAI format on port twelve thirty-four. There is also a command-line tool, l m s, for scripting. On Apple silicon it runs both GGUF and MLX models. It is perfect for letting a marketing or operations team try local models on real work before anyone buys a server.
Worked example: UK charity
A real-style example. A UK charity handles safeguarding notes that must never leave the device. Case workers use LM Studio on their laptops with one small approved model and an IT setup guide. The data team runs llama server on an internal Linux box, with an API key, firewall rules and a JSON schema that extracts risk categories straight into their case system. IT records both in the model register and pins model files by checksum. Convenience for people, control for pipelines.
Simple example: bookshop email tags
A simple example of constrained output. A small online bookshop wants to tag incoming emails with one of four labels: order status, return, recommendation request or other. Without constraints, the model sometimes replies with a friendly sentence instead of a label. With a JSON schema that allows only those four strings, every single reply is one of the four. The shop's automation never breaks on a surprise sentence again. They still check a sample each week to make sure the labels are right, not just valid.
Which one when?
How do you decide between them for a given team? Ask who will operate it and what it feeds. If a person is chatting and exploring, LM Studio's interface and fit-aware downloads save hours. If software depends on the output, a pipeline, a database, an integration, you want the llama dot cpp server, or a bigger engine, with fixed flags, an API key and a schema. Many organizations use both, with the same approved model files, recorded in the same register.
FAQ
A common question: if Ollama and LM Studio both use llama dot cpp underneath, why would I ever use it directly? Three reasons. You get new features and model support first, sometimes weeks earlier. You get precise control over flags like context per slot, cache types and grammars. And you get a tiny, dependency-light server that is easy to put in a container on almost any hardware. Another question: are GGUF files from LM Studio and Ollama interchangeable? The GGUF format is shared, so a GGUF file generally works in any llama dot cpp based runtime, though each app manages its own storage and naming.
Try this now
Try this now. Start llama server with an API key and a small model, and send it one request with a JSON schema for a task you care about. Then send the same request without the schema, ten times, and count how many answers you could parse. That tiny experiment will convince your team faster than any slide.
Watch me do it
Watch me do it. I start llama server with a GGUF model, a context of sixteen thousand, all layers on the GPU, two parallel slots, the jinja template flag and an API key from an environment variable. I check the health endpoint: it says ok. Now I send a classification request for an email that says my order arrived broken, I want my money back, with a JSON schema that allows only four intents. The response comes back as clean JSON: intent refund, order number included. I send the same request ten times without the schema and count: eight parseable, two chatty answers. Next I switch to LM Studio for a colleague. I search the same model, and the app shows which quantization fits her laptop. She downloads it, compares it side by side with another model on three of her real emails, then turns on the local server. Two tools, two audiences, same model file.
Recap
Recap. llama dot cpp gives you control: flags, schemas, performance tuning and a small OpenAI-compatible server. Remember context is split across slots, and templates matter for tools. LM Studio gives people a friendly way to explore and serve models locally. Your next step: serve a model with an API key and a JSON schema for one classification task from your work, then benchmark two quantizations with llama bench.
Key takeaways
- llama.cpp gives maximum control and a lightweight OpenAI-compatible server
- Context (-c) is shared across --parallel slots
- Grammar/JSON-schema constraints guarantee format validity
- LM Studio is a friendly desktop app with a local server on port 1234 by default
- Benchmark with llama-bench on your own hardware
Try it
Serve a GGUF with llama-server using an API key and a JSON schema for a classification task from your work; benchmark two quantizations with llama-bench.