Open-Weight and Local AI: Run, Choose and Deploy Your Own ModelsRunning models locally · Lesson 7 of 16
Ollama hands-on: models, Modelfiles and the API
Video lecture
Ollama hands-on: models, Modelfiles and the API
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Ollama hands-on
In the next ten minutes you will go from an empty laptop to a branded, team-ready AI assistant that never sends a word to the cloud. We will use Ollama, the most popular local model runner. You will learn the five commands that matter, the one setting that silently breaks long documents, how to package your own assistant, how to call it from code, and how to share it without accidentally publishing it to the internet.
0:33 Analogy: app store + power socket
Think of Ollama as an app store and a power socket in one. The library is the app store: you browse, pull, and the model lands on your disk. The local server is the socket: once a model is plugged in, any app on your machine can draw from it through a standard plug, the API. And a Modelfile is like a preset on a kitchen appliance. You set the temperature, the timer and the recipe once, save it with a name, and everyone in the house gets the same result every time.
1:13 Five commands
The five commands. Ollama pull downloads a model. Ollama run starts a chat. Ollama list shows what you have and how big it is. Ollama show tells you the details, including quantization, template and license. And my favorite diagnostic, ollama ps, shows what is loaded and whether it sits fully on your GPU. If you see part of it on the CPU, it did not fit, and it will be slower. Pick a smaller model, a smaller quantization, or a shorter context.
1:49 The context gotcha
Now the gotcha that catches almost everyone. A model might support a very long context, but the runtime gives it a smaller default to save memory. So you paste a twenty-page contract, and the model cheerfully answers as if page one never existed. It was truncated. The fix is to set the context length on purpose, per request, in a Modelfile, or server-wide. Just remember from the sizing lesson: more context means more memory.
2:21 Modelfile
A Modelfile packages your assistant. It names a base model, sets parameters like temperature and context length, and adds a system prompt. For a fictional Dubai travel agency, the system prompt says: answer in the language of the question, English or Arabic, say when you are unsure, and never invent prices or visa rules. One command, ollama create, turns it into a named model everyone on the team runs identically. You can also point a Modelfile at a GGUF file you built yourself.
2:57 Calling it from code
To use it from software, Ollama listens on port one one four three four on your machine. There is a native API with chat and embedding endpoints, and an OpenAI-compatible API under slash v one. That second one is the quiet superpower. Code written for the OpenAI software kit can talk to your local model by changing two things: the base address and the model name. Later in the course, that exact trick is how we route between local and cloud models.
3:33 Share it safely
Now safety. By default, Ollama only listens to your own machine. You can open it to the network, but its API has no built-in login. Security researchers regularly find exposed model servers on the internet. So if you share it, put it behind a reverse proxy that requires authentication and encryption, restrict it to your office network or VPN, and never expose the raw port publicly. Also pin a specific model tag rather than latest, so an update never surprises you.
4:08 Worked example: travel agency
Here is the whole thing working. The fictional travel agency installs Ollama on a desktop with a sixteen gigabyte GPU, creates its assistant with a sixteen thousand token context, and puts it behind the office VPN with a password-protected proxy. Eight agents use a simple web chat that talks to the OpenAI-compatible endpoint. Passport scans never leave the building. When a question needs current visa rules, the assistant says to check the official portal.
4:40 Simple example: quiz maker
A simple example. A teacher in Leeds wants a private assistant that turns lesson notes into five quiz questions. She installs Ollama, pulls a small model, and writes a four-line Modelfile: the base model, a low temperature, an eight thousand token context, and a system prompt saying always return five multiple-choice questions with answers. She runs ollama create quiz maker, and from then on, one command turns any lesson note into a quiz. No accounts, no uploads, no student data leaving her laptop.
5:16 Two habits
Two small habits make Ollama deployments far easier to support. First, keep your Modelfiles in version control alongside the application code, so a change to the system prompt or context length is reviewed like any other change. Second, keep a short runbook: which model tag is approved, how to update it, how to check ollama ps, how to restart the service, and who to call. When the person who set it up goes on holiday, the runbook is what keeps the assistant running.
5:52 FAQ
Two questions people ask after their first day with Ollama. First: why is the first answer slow and the rest fast? Because the model loads from disk into memory on the first request. Ollama keeps it loaded for a while afterwards, and you can adjust how long with its keep alive setting, so during office hours you can keep it warm. Second: can I use my own fine-tuned model? Yes. If you have a GGUF file, a Modelfile that points FROM that file turns it into a named model, and Ollama can also attach supported adapters. Just record it in your model register like any other model.
6:39 Try this now
Try this now. Install Ollama, pull one small model, and run ollama ps while it is answering a question. Look at the processor column. If it says one hundred percent GPU, great. If it shows a CPU share, try a smaller model or a shorter context and watch the change. Then write a Modelfile for a task you do every week, give it a name, and use it for the rest of the week.
7:11 Watch me do it
Watch me do it. I open a terminal and type ollama pull, then a small model tag. While it downloads, I create a file called Modelfile. First line: FROM and the model tag. Then PARAMETER temperature zero point two, and PARAMETER num context sixteen thousand three hundred eighty-four. Then a SYSTEM block: you are the internal assistant for our support team; answer in the language of the question; say when you are unsure. I save it and run ollama create support assistant with the file. Now ollama run support assistant, and I paste a long refund policy plus a question about the last paragraph. It answers correctly, which tells me the context setting worked. In a second terminal I run ollama ps: one hundred percent GPU. Finally I call it from Python with the OpenAI client, changing only the base address and the model name. Ten minutes, start to finish.
8:16 Recap
Recap. Ollama pulls, manages and serves models locally. Check ollama ps for GPU fit. Set context length deliberately. Package assistants with Modelfiles. Use the OpenAI-compatible endpoint to reuse code. And never expose it without authentication. Your next step: write a Modelfile for an assistant in your own domain and call it from Python with both clients. The code is in the lesson.
What Ollama is
Ollama is a local model runner for macOS, Windows and Linux. It downloads models from its library (mostly GGUF under the hood), manages them on disk, loads them onto your GPU or CPU, and exposes a local HTTP API, including an OpenAI-compatible endpoint. For most individuals and small teams it is the fastest way from "nothing" to "working local model".
Core commands
ollama pull gpt-oss:20b # download a model (tags change; browse the Ollama library)
ollama run gpt-oss:20b # interactive chat in the terminal
ollama list # installed models and their sizes
ollama ps # loaded models, memory used, CPU/GPU split
ollama show gpt-oss:20b # details: parameters, quantization, template, license
ollama rm old-model:tag # free disk spaceollama ps is your first diagnostic. If it shows a model partly on CPU, it did not fit fully in GPU memory and will be slower; choose a smaller model or quantization, or reduce context.
The context-length gotcha
A model may support a very long context, but the runtime allocates a smaller default to save memory. If you paste a long document and the model "forgets" the beginning, the input was probably truncated. Set the context explicitly, per request (num_ctx option), in a Modelfile (PARAMETER num_ctx), or server-wide via the environment variable documented for your Ollama version (for example OLLAMA_CONTEXT_LENGTH). Remember the previous module: more context means more KV cache memory.
Modelfiles: package your own assistant
A Modelfile bundles a base model with parameters and a system prompt, so everyone on the team runs the same configuration.
# Modelfile
FROM qwen3:8b
PARAMETER temperature 0.2
PARAMETER num_ctx 16384
SYSTEM """You are the internal assistant for Crescent Travel (Dubai).
Answer in the language of the question (English or Arabic).
If you are not sure, say so and suggest who to ask. Never invent prices or visa rules."""ollama create crescent-assistant -f Modelfile
ollama run crescent-assistant "What should I check before booking a Schengen visa appointment?"You can also FROM ./path/to/model.gguf to import a GGUF you built or downloaded yourself.
The native API
Ollama listens on http://localhost:11434 by default.
curl http://localhost:11434/api/chat -d '{
"model": "crescent-assistant",
"messages": [{"role": "user", "content": "Summarize our refund rules in 3 bullets."}],
"stream": false,
"options": {"num_ctx": 16384, "temperature": 0.2}
}'Embeddings (for RAG, Module 6) use /api/embed with an embedding model such as nomic-embed-text or another embedding model from the library.
Python: native client and OpenAI-compatible client
# pip install ollama
from ollama import chat, ResponseError
try:
resp = chat(model="crescent-assistant",
messages=[{"role": "user", "content": "Give me a 3-step checklist for a UAE tourist visa inquiry."}],
options={"temperature": 0.2})
print(resp.message.content)
except ResponseError as e:
print("Ollama error:", e.error) # e.g. model not found: run `ollama pull` firstBecause Ollama also serves an OpenAI-compatible API at /v1, code written for the OpenAI SDK can point at it by changing the base URL. The API key is required by the SDK but ignored by Ollama:
# pip install openai
import os
from openai import OpenAI
client = OpenAI(base_url=os.getenv("LLM_BASE_URL", "http://localhost:11434/v1"),
api_key=os.getenv("LLM_API_KEY", "ollama"))
r = client.chat.completions.create(model="crescent-assistant",
messages=[{"role": "user", "content": "Hello in Arabic and English"}])
print(r.choices[0].message.content)This is the foundation of hybrid routing later: the same code can talk to a local model or a cloud API by switching LLM_BASE_URL and the model name.
Serving to your team (safely)
By default Ollama binds to 127.0.0.1, so only your machine can reach it. To share it on a LAN you can set OLLAMA_HOST=0.0.0.0, but Ollama's API has no built-in authentication. Internet scans regularly find exposed model servers. If you share it:
- Put it behind a reverse proxy (for example Caddy or Nginx) that enforces authentication and TLS.
- Restrict by firewall or VPN to your office network.
- Never expose port 11434 directly to the internet.
Worked example: a Dubai travel agency
Crescent Travel (fictional) has eight agents who answer visa and package questions. The owner installs Ollama on a desktop with a 16 GB GPU, creates the crescent-assistant Modelfile with a 16k context, and puts it behind the office VPN with a Caddy proxy requiring a password. Agents use a simple web chat front end pointing at the OpenAI-compatible endpoint. Customer passport scans never leave the office. When questions need current visa rules, agents check the official government portal; the assistant is told to say so.
Pitfalls
- Silent truncation from default context size.
- Exposing the API without authentication.
- Using "latest" tags in production. Pin a specific tag and record it in your model register.
- Assuming the Modelfile system prompt is a security boundary; users can still try to override it.
How to measure success
Your team can run a pinned, documented assistant configuration; ollama ps shows it fully on GPU at your chosen context; and the API is reachable only through an authenticated path.
Key takeaways
- Ollama pulls, manages and serves local models with a native and an OpenAI-compatible API
- Check ollama ps: CPU offload means it did not fit in GPU memory
- Set context length explicitly to avoid silent truncation
- Modelfiles package base model, parameters and system prompt for consistent team use
- Never expose the API without authentication; pin model tags
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Create a Modelfile for an assistant in your domain, set num_ctx deliberately, and call it from Python with both the native and the OpenAI-compatible client.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.