Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsRunning models locally · Lesson 7 of 16

Ollama hands-on: models, Modelfiles and the API

Article · 16 min · 9 min lecture

Video lecture

Ollama hands-on: models, Modelfiles and the API

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

Ollama hands-on

  • Five commands
  • The context-length gotcha
  • Modelfiles, API, safe sharing

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What Ollama is

Ollama is a local model runner for macOS, Windows and Linux. It downloads models from its library (mostly GGUF under the hood), manages them on disk, loads them onto your GPU or CPU, and exposes a local HTTP API, including an OpenAI-compatible endpoint. For most individuals and small teams it is the fastest way from "nothing" to "working local model".

Core commands

ollama pull gpt-oss:20b        # download a model (tags change; browse the Ollama library)
ollama run gpt-oss:20b         # interactive chat in the terminal
ollama list                    # installed models and their sizes
ollama ps                      # loaded models, memory used, CPU/GPU split
ollama show gpt-oss:20b        # details: parameters, quantization, template, license
ollama rm old-model:tag        # free disk space

ollama ps is your first diagnostic. If it shows a model partly on CPU, it did not fit fully in GPU memory and will be slower; choose a smaller model or quantization, or reduce context.

The context-length gotcha

A model may support a very long context, but the runtime allocates a smaller default to save memory. If you paste a long document and the model "forgets" the beginning, the input was probably truncated. Set the context explicitly, per request (num_ctx option), in a Modelfile (PARAMETER num_ctx), or server-wide via the environment variable documented for your Ollama version (for example OLLAMA_CONTEXT_LENGTH). Remember the previous module: more context means more KV cache memory.

Modelfiles: package your own assistant

A Modelfile bundles a base model with parameters and a system prompt, so everyone on the team runs the same configuration.

# Modelfile
FROM qwen3:8b
PARAMETER temperature 0.2
PARAMETER num_ctx 16384
SYSTEM """You are the internal assistant for Crescent Travel (Dubai).
Answer in the language of the question (English or Arabic).
If you are not sure, say so and suggest who to ask. Never invent prices or visa rules."""
ollama create crescent-assistant -f Modelfile
ollama run crescent-assistant "What should I check before booking a Schengen visa appointment?"

You can also FROM ./path/to/model.gguf to import a GGUF you built or downloaded yourself.

The native API

Ollama listens on http://localhost:11434 by default.

curl http://localhost:11434/api/chat -d '{
  "model": "crescent-assistant",
  "messages": [{"role": "user", "content": "Summarize our refund rules in 3 bullets."}],
  "stream": false,
  "options": {"num_ctx": 16384, "temperature": 0.2}
}'

Embeddings (for RAG, Module 6) use /api/embed with an embedding model such as nomic-embed-text or another embedding model from the library.

Python: native client and OpenAI-compatible client

# pip install ollama
from ollama import chat, ResponseError

try:
    resp = chat(model="crescent-assistant",
                messages=[{"role": "user", "content": "Give me a 3-step checklist for a UAE tourist visa inquiry."}],
                options={"temperature": 0.2})
    print(resp.message.content)
except ResponseError as e:
    print("Ollama error:", e.error)   # e.g. model not found: run `ollama pull` first

Because Ollama also serves an OpenAI-compatible API at /v1, code written for the OpenAI SDK can point at it by changing the base URL. The API key is required by the SDK but ignored by Ollama:

# pip install openai
import os
from openai import OpenAI

client = OpenAI(base_url=os.getenv("LLM_BASE_URL", "http://localhost:11434/v1"),
                api_key=os.getenv("LLM_API_KEY", "ollama"))
r = client.chat.completions.create(model="crescent-assistant",
                                   messages=[{"role": "user", "content": "Hello in Arabic and English"}])
print(r.choices[0].message.content)

This is the foundation of hybrid routing later: the same code can talk to a local model or a cloud API by switching LLM_BASE_URL and the model name.

Serving to your team (safely)

By default Ollama binds to 127.0.0.1, so only your machine can reach it. To share it on a LAN you can set OLLAMA_HOST=0.0.0.0, but Ollama's API has no built-in authentication. Internet scans regularly find exposed model servers. If you share it:

  • Put it behind a reverse proxy (for example Caddy or Nginx) that enforces authentication and TLS.
  • Restrict by firewall or VPN to your office network.
  • Never expose port 11434 directly to the internet.

Worked example: a Dubai travel agency

Crescent Travel (fictional) has eight agents who answer visa and package questions. The owner installs Ollama on a desktop with a 16 GB GPU, creates the crescent-assistant Modelfile with a 16k context, and puts it behind the office VPN with a Caddy proxy requiring a password. Agents use a simple web chat front end pointing at the OpenAI-compatible endpoint. Customer passport scans never leave the office. When questions need current visa rules, agents check the official government portal; the assistant is told to say so.

Pitfalls

  • Silent truncation from default context size.
  • Exposing the API without authentication.
  • Using "latest" tags in production. Pin a specific tag and record it in your model register.
  • Assuming the Modelfile system prompt is a security boundary; users can still try to override it.

How to measure success

Your team can run a pinned, documented assistant configuration; ollama ps shows it fully on GPU at your chosen context; and the API is reachable only through an authenticated path.

Key takeaways

  • Ollama pulls, manages and serves local models with a native and an OpenAI-compatible API
  • Check ollama ps: CPU offload means it did not fit in GPU memory
  • Set context length explicitly to avoid silent truncation
  • Modelfiles package base model, parameters and system prompt for consistent team use
  • Never expose the API without authentication; pin model tags

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An agent pastes a 20-page contract and the model ignores the first pages. What is the most likely cause?
  2. You want colleagues on the office network to use your Ollama server. What is the safest approach?
  3. Why is the OpenAI-compatible endpoint useful?

Put it into practice

Create a Modelfile for an assistant in your domain, set num_ctx deliberately, and call it from Python with both the native and the OpenAI-compatible client.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.