Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsRunning models locally · Lesson 9 of 16

Local models for real work: structured output, tools and reliability

Article · 15 min · 8 min lecture

Video lecture

Local models for real work: structured output, tools and reliability

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Local models for real work

  • Structured output
  • Tool calling
  • Reliability patterns

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

From chat toy to system component

A local model becomes useful to a business when other software can depend on it: extract fields into a CRM, classify tickets, call tools, and fail predictably. This lesson turns your local runtime into a reliable component.

Structured output three ways

  1. Constrained decoding (best). The runtime enforces a JSON schema (Ollama's format parameter, llama.cpp schemas/grammars, vLLM and SGLang guided decoding). Output always parses.
  2. JSON mode. The runtime ensures valid JSON but not your specific schema.
  3. Prompt-only. You ask nicely for JSON. Works most of the time; fails at the worst time.

Use (1) whenever your runtime supports it, and still validate on your side.

Ollama structured output with Pydantic

# pip install ollama pydantic
from typing import Literal, Optional
from pydantic import BaseModel, ValidationError
from ollama import chat

class Lead(BaseModel):
    name: str
    company: Optional[str] = None
    country: Optional[str] = None
    budget_usd: Optional[int] = None
    intent: Literal["demo", "pricing", "support", "partnership", "other"]

email = """Salam, I'm Ayesha from Indus Textiles in Faisalabad. We'd like pricing
for 25 seats and maybe a demo next week. Budget is around 3,000 USD per year."""

resp = chat(model="qwen3:8b",
            messages=[{"role": "system", "content": "Extract lead details. Use null when unknown."},
                      {"role": "user", "content": email}],
            format=Lead.model_json_schema(),   # constrained to this schema
            options={"temperature": 0})
try:
    lead = Lead.model_validate_json(resp.message.content)
    print(lead)
except ValidationError as e:
    print("Validation failed, route to human:", e)

The schema guarantees shape; your validation and evaluation guarantee meaning. Note the ambiguity in the email: "pricing ... and maybe a demo". Decide in your labeling guide which intent wins, then test for it.

Tool calling locally

Many open-weight models are trained for tool (function) calling, and Ollama, llama.cpp (with the right template), LM Studio, vLLM and SGLang expose it through OpenAI-style tools. The loop is the same as with cloud APIs:

from ollama import chat

def get_order_status(order_id: str) -> str:
    """Look up an order's delivery status.

    Args:
      order_id: The order number, e.g. "5531"
    """
    fake_db = {"5531": "Out for delivery in Jeddah"}   # replace with your real API call
    return fake_db.get(order_id, "Order not found")

messages = [{"role": "user", "content": "Where is order 5531?"}]
resp = chat(model="qwen3:8b", messages=messages, tools=[get_order_status])

if resp.message.tool_calls:
    messages.append(resp.message)
    for call in resp.message.tool_calls:
        if call.function.name == "get_order_status":
            result = get_order_status(**call.function.arguments)
            messages.append({"role": "tool", "content": result, "tool_name": call.function.name})
    final = chat(model="qwen3:8b", messages=messages)
    print(final.message.content)
else:
    print(resp.message.content)

The Ollama Python library can build the tool schema from a function's signature and docstring. Always check the tool name against an allow-list and validate arguments before executing anything.

Reliability patterns

  • Temperature 0 for extraction and classification; higher only for creative drafting.
  • Retries with repair. If validation fails, retry once with the validation error in the prompt; after that, route to a human queue.
  • Timeouts. Local models can stall under load; set client timeouts and a fallback path.
  • Warm models. The first request after idle loads the model from disk. Keep models loaded during business hours (Ollama's keep_alive setting) to avoid cold-start latency.
  • Version pinning. Log the model tag and quantization with every output so you can trace regressions.
  • Prompt injection still applies. A local model reading an email can still be manipulated by text in that email. Keep tools least-privilege and require confirmation for side effects.

Worked example: Jeddah e-commerce support triage

An online store in Jeddah receives Arabic and English emails. A local 8B model on an office GPU extracts intent, order_id and language with a schema, then a tool call fetches order status from the store's API (read-only). Refund intents are never executed automatically: they create a draft for a human agent. After two weeks the team measures (illustrative): 94% of intents correct on a 300-email labeled sample, zero malformed outputs, median latency 1.4 s.

Hands-on checklist

  1. Define a Pydantic schema for one extraction task.
  2. Label 50 real examples by hand (the "gold set").
  3. Run the local model with the schema at temperature 0.
  4. Compute accuracy per field; list the 10 worst errors.
  5. Fix the labeling guide or prompt, rerun, and compare.

Pitfalls

  • Trusting schema-valid output as correct output.
  • Letting a model call write-actions without confirmation.
  • Measuring only on English when customers write Arabic or Urdu.
  • Forgetting cold-start latency in user-facing flows.

How to measure success

Per-field accuracy on a labeled gold set, zero parse failures, a p95 latency you can live with, and a documented human fallback path.

Key takeaways

  • Prefer constrained decoding (schemas) over prompt-only JSON
  • Schema-valid is not the same as correct; validate and evaluate on a gold set
  • Tool calling works locally with the same loop as cloud APIs; allow-list tools
  • Use temperature 0, retries with repair, timeouts and warm models
  • Prompt injection and least-privilege still apply locally

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which approach gives the strongest guarantee that output parses as your schema?
  2. A local model reads customer emails and can call a refund tool. What is the safest design?
  3. The first request each morning is slow, then speed is fine. What is likely happening?

Put it into practice

Build a schema-constrained extractor for one task in your work, label 50 examples, and report per-field accuracy and the 10 worst errors.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.