Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsRunning models locally · Lesson 9 of 16
Local models for real work: structured output, tools and reliability
Video lecture
Local models for real work: structured output, tools and reliability
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Local models for real work
A chatbot that is right most of the time is a demo. A component that is right, parseable and traceable every time is a product. In this lesson you will turn a local model into something other software can rely on: guaranteed structured output, tool calling, and the reliability patterns that stop two a.m. pages. Everything runs locally, and the code is in the lesson.
0:28 Analogy: paper form vs web form
Here is an analogy for structured output. Asking a model for JSON in plain words is like asking a new employee to fill in a paper form by hand. Most of the time it is fine, but sometimes a box is skipped or someone writes outside the lines. A schema with constrained decoding is a web form with dropdowns and required fields. The person can only submit valid entries. It does not guarantee they chose the right option, but it guarantees the form is complete and machine-readable.
1:06 Three levels of structure
There are three levels of structured output. Level one, prompt-only: you ask nicely for JSON, and it works until it does not. Level two, JSON mode: valid JSON, but not necessarily your fields. Level three, constrained decoding: you hand the runtime a schema and it only allows tokens that fit. Ollama does this with its format parameter, llama dot cpp with schemas and grammars, vLLM and SGLang with guided decoding. Use level three whenever you can, and still validate on your side.
1:42 Pydantic + Ollama format
Here is the pattern in Python. You describe the lead you want with a Pydantic model: name, company, country, budget, and an intent that must be demo, pricing, support, partnership or other. You pass the schema to Ollama, set temperature to zero, then validate the result. Try it on an email like this: Ayesha from a textile firm in Faisalabad wants pricing for twenty-five seats and maybe a demo. Notice the ambiguity. Pricing or demo? The schema cannot decide that. Your labeling guide must, and your tests must check it.
2:21 Tool-calling loop
Tool calling works locally with the same loop you know from cloud APIs. The model sees tool descriptions, decides to call one, say get order status with order five five three one. Your code checks the tool name against an allow-list, validates the arguments, runs the function, and sends the result back. Then the model writes the final answer. The Ollama Python library can even build the tool description from your function's signature and docstring. Just make sure the model's chat template supports tools.
2:58 Reliability patterns
Now reliability. Use temperature zero for extraction and classification. If validation fails, retry once with the error message, then route to a human. Set client timeouts, because a busy local server can stall. Keep the model warm during business hours, since the first request after idle has to load it from disk. Log the model tag and quantization with every output so you can trace a regression. And remember, prompt injection works locally too. An email can contain instructions. Keep tools read-only where possible and confirm anything with side effects.
3:37 Worked example: Jeddah store (illustrative)
A worked example. An online store in Jeddah gets Arabic and English emails. A local eight billion model on an office GPU extracts intent, order number and language with a schema. A read-only tool fetches order status. Refunds are never executed automatically; they become drafts for a human agent. After two weeks, on three hundred labeled emails, the team sees, and these are illustrative, ninety-four percent correct intents, zero malformed outputs, and a median latency around one and a half seconds.
4:12 Hands-on: the gold-set loop
Your hands-on exercise is a gold-set loop. Define a schema for one extraction task. Label fifty real examples by hand. Run the model at temperature zero. Compute accuracy per field and read the ten worst errors. Fix your labeling guide or prompt, run again, compare. That loop, not the model name, is what makes local AI trustworthy.
4:37 Simple example: studio timetable
A simple example of tool calling. A fitness studio in Riyadh wants its local assistant to answer, what time is the next yoga class? The model does not know the timetable, so you give it one read-only tool, get next class, with a class type argument. The model calls get next class with yoga, your code looks it up in the booking system, returns seven p m, and the model replies: the next yoga class is at seven this evening. The model never touches the booking system directly. Your code does, with only the permission to read.
5:19 Valid but wrong
Let me show you a failure that schemas do not catch. A support email says: I was charged twice, and the parcel never arrived. Is that a refund intent or a delivery intent? The model returns perfectly valid JSON with intent set to delivery. Parse succeeds, accuracy fails. The fix is not a better model. It is a labeling rule, for example, money issues take priority, plus a test case for exactly this pattern in your gold set. Most accuracy gains in production come from sharpening these rules.
5:57 FAQ: are small models good at tools?
A question worth answering: do small local models handle tool calling reliably? Increasingly yes, but it varies by model and by template. The practical approach is to test tool calling as part of your evaluation, with ten or twenty requests that should call a tool and ten that should not. Measure both: did it call the right tool with valid arguments, and did it avoid calling tools when it should just answer? If a model struggles, give it fewer, clearer tools with good descriptions, or route tool-heavy requests to a stronger model.
6:37 Try this now
Try this now. Take one form you fill in by hand at work, maybe a lead intake or an expense note, and turn its fields into a Pydantic model. Then feed the local model three real examples with the schema and temperature zero. Count how many fields are right. That is your first gold-set measurement, and it takes about twenty minutes.
7:04 Watch me do it
Watch me do it. I write a Pydantic class for a sales lead: name, company, country, budget and an intent from a fixed list. I pass its schema to Ollama's format parameter with temperature zero, and send three real, anonymized emails. The first comes back perfect. The second has budget empty, correct, because the email did not mention one. The third labels intent as demo, but my labeling guide says pricing wins when both are mentioned. So I add one sentence to the system prompt about that priority, rerun, and it now says pricing. Next I add a read-only tool, get order status, and ask where is order five five three one. The model proposes the call, my code checks the name against an allow-list, runs it, and returns the result, and the model answers in one sentence. I log the model tag and quantization with every output. Twenty minutes, and it behaves like a component.
8:12 Recap
Recap. Constrain output with schemas, but remember valid is not the same as correct. Tool calling works locally with an allow-list and validation. Build reliability with low temperature, retries, timeouts, warm models, logging and least privilege. Your next step: build one schema-constrained extractor for your work and measure it on fifty labeled examples.
From chat toy to system component
A local model becomes useful to a business when other software can depend on it: extract fields into a CRM, classify tickets, call tools, and fail predictably. This lesson turns your local runtime into a reliable component.
Structured output three ways
- Constrained decoding (best). The runtime enforces a JSON schema (Ollama's
formatparameter, llama.cpp schemas/grammars, vLLM and SGLang guided decoding). Output always parses. - JSON mode. The runtime ensures valid JSON but not your specific schema.
- Prompt-only. You ask nicely for JSON. Works most of the time; fails at the worst time.
Use (1) whenever your runtime supports it, and still validate on your side.
Ollama structured output with Pydantic
# pip install ollama pydantic
from typing import Literal, Optional
from pydantic import BaseModel, ValidationError
from ollama import chat
class Lead(BaseModel):
name: str
company: Optional[str] = None
country: Optional[str] = None
budget_usd: Optional[int] = None
intent: Literal["demo", "pricing", "support", "partnership", "other"]
email = """Salam, I'm Ayesha from Indus Textiles in Faisalabad. We'd like pricing
for 25 seats and maybe a demo next week. Budget is around 3,000 USD per year."""
resp = chat(model="qwen3:8b",
messages=[{"role": "system", "content": "Extract lead details. Use null when unknown."},
{"role": "user", "content": email}],
format=Lead.model_json_schema(), # constrained to this schema
options={"temperature": 0})
try:
lead = Lead.model_validate_json(resp.message.content)
print(lead)
except ValidationError as e:
print("Validation failed, route to human:", e)The schema guarantees shape; your validation and evaluation guarantee meaning. Note the ambiguity in the email: "pricing ... and maybe a demo". Decide in your labeling guide which intent wins, then test for it.
Tool calling locally
Many open-weight models are trained for tool (function) calling, and Ollama, llama.cpp (with the right template), LM Studio, vLLM and SGLang expose it through OpenAI-style tools. The loop is the same as with cloud APIs:
from ollama import chat
def get_order_status(order_id: str) -> str:
"""Look up an order's delivery status.
Args:
order_id: The order number, e.g. "5531"
"""
fake_db = {"5531": "Out for delivery in Jeddah"} # replace with your real API call
return fake_db.get(order_id, "Order not found")
messages = [{"role": "user", "content": "Where is order 5531?"}]
resp = chat(model="qwen3:8b", messages=messages, tools=[get_order_status])
if resp.message.tool_calls:
messages.append(resp.message)
for call in resp.message.tool_calls:
if call.function.name == "get_order_status":
result = get_order_status(**call.function.arguments)
messages.append({"role": "tool", "content": result, "tool_name": call.function.name})
final = chat(model="qwen3:8b", messages=messages)
print(final.message.content)
else:
print(resp.message.content)The Ollama Python library can build the tool schema from a function's signature and docstring. Always check the tool name against an allow-list and validate arguments before executing anything.
Reliability patterns
- Temperature 0 for extraction and classification; higher only for creative drafting.
- Retries with repair. If validation fails, retry once with the validation error in the prompt; after that, route to a human queue.
- Timeouts. Local models can stall under load; set client timeouts and a fallback path.
- Warm models. The first request after idle loads the model from disk. Keep models loaded during business hours (Ollama's
keep_alivesetting) to avoid cold-start latency. - Version pinning. Log the model tag and quantization with every output so you can trace regressions.
- Prompt injection still applies. A local model reading an email can still be manipulated by text in that email. Keep tools least-privilege and require confirmation for side effects.
Worked example: Jeddah e-commerce support triage
An online store in Jeddah receives Arabic and English emails. A local 8B model on an office GPU extracts intent, order_id and language with a schema, then a tool call fetches order status from the store's API (read-only). Refund intents are never executed automatically: they create a draft for a human agent. After two weeks the team measures (illustrative): 94% of intents correct on a 300-email labeled sample, zero malformed outputs, median latency 1.4 s.
Hands-on checklist
- Define a Pydantic schema for one extraction task.
- Label 50 real examples by hand (the "gold set").
- Run the local model with the schema at temperature 0.
- Compute accuracy per field; list the 10 worst errors.
- Fix the labeling guide or prompt, rerun, and compare.
Pitfalls
- Trusting schema-valid output as correct output.
- Letting a model call write-actions without confirmation.
- Measuring only on English when customers write Arabic or Urdu.
- Forgetting cold-start latency in user-facing flows.
How to measure success
Per-field accuracy on a labeled gold set, zero parse failures, a p95 latency you can live with, and a documented human fallback path.
Key takeaways
- Prefer constrained decoding (schemas) over prompt-only JSON
- Schema-valid is not the same as correct; validate and evaluate on a gold set
- Tool calling works locally with the same loop as cloud APIs; allow-list tools
- Use temperature 0, retries with repair, timeouts and warm models
- Prompt injection and least-privilege still apply locally
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Build a schema-constrained extractor for one task in your work, label 50 examples, and report per-field accuracy and the 10 worst errors.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.