Skip to content

Open-Weight and Local AI: Run, Choose and Deploy Your Own Models · The open-weight landscape · lesson 3 of 16 · 15 min

Model families in 2026 and how to shortlist

The landscape changes monthly; your method should not

New open-weight releases land almost every week. Instead of memorizing a leaderboard, learn the dimensions that decide fit, know the major families at a high level, and keep a repeatable shortlist process. Treat every specific below as "true as of September 2026; check the model card".

The dimensions that matter

  • Size (parameters). Roughly: 1–4B for phones and simple tasks; 7–15B for laptops and many business tasks; 20–40B for workstation-class quality; 100B+ for server-grade quality.
  • Dense vs Mixture-of-Experts (MoE). A dense model uses every parameter for every token. An MoE model has many "experts" and routes each token through a few, so it has total parameters (what must fit in memory) and active parameters (what drives compute per token). gpt-oss-120b, for example, has about 117B total and 5.1B active parameters per its model card: it needs lots of memory but generates quickly for its size.
  • Reasoning mode. Many families ship "thinking" variants or a switchable reasoning effort. Better on multi-step problems, slower and more tokens.
  • Context length. Advertised maxima are large; quality at long context varies and memory cost grows with context (next module).
  • Modalities. Text only, or also images, audio, video.
  • Languages. If you serve Arabic, Urdu or code-mixed "Roman Urdu" users, test them explicitly.
  • Tool calling and structured output. Essential for agents; quality varies by family and by the chat template your runtime uses.
  • License. Previous lesson.

The major families (September 2026 snapshot)

| Family | Publisher | What it is known for | License (verify per release) | |---|---|---|---| | gpt-oss (20b, 120b) | OpenAI | MoE reasoning models with adjustable reasoning effort, tool use; shipped natively quantized in MXFP4 so 120b fits a single 80 GB GPU and 20b runs within 16 GB | Apache-2.0 | | Gemma 4 | Google | Range from edge-sized models to ~31B; strong quality per parameter; multimodal variants | Apache-2.0 (new with Gemma 4) | | Qwen3 / Qwen3.5 | Alibaba | Very broad size range from small dense to large MoE; strong multilingual and coding variants | Mostly Apache-2.0 | | DeepSeek (R1, V3.x, V4) | DeepSeek | Very large MoE models with strong reasoning; distilled smaller variants exist | MIT | | Mistral 3 (Large 3, Ministral 3) | Mistral AI | Large 3 is a sparse MoE; Ministral 3 dense 3B/8B/14B with base, instruct and reasoning variants for edge | Apache-2.0 | | Llama 4 (Scout, Maverick) | Meta | MoE multimodal models with long context | Llama 4 Community License | | Phi-4 family | Microsoft | Small models (mini, reasoning, multimodal) strong for their size | MIT |

Other notable publishers release strong open weights too (for example NVIDIA's Nemotron, Zhipu's GLM, Moonshot's Kimi, IBM's Granite, Ai2's fully open OLMo). Meta has signaled that some future frontier models may launch closed first, which is a reminder that a family's openness can change.

A four-step shortlist process

  1. Filter by hard constraints: license, languages, modalities, hardware budget (memory you can afford), context you really need.
  2. Pick 3–5 candidates spanning sizes: one small, one mid, one large.
  3. Run your own evaluation set (30–100 real tasks with expected answers or rubrics) at the quantization you will deploy.
  4. Choose the smallest model that passes, then measure latency and cost on target hardware.

"Smallest that passes" is the key discipline. Every step down in size cuts memory, cost and latency.

Worked example: an agency in Manchester

A content agency wants local drafting of social posts in English and Arabic for Gulf clients, plus JSON-formatted campaign briefs.

  • Constraints: Apache or MIT preferred; Arabic essential; 24 GB GPU per workstation; 16k context is plenty.
  • Candidates: a Qwen3.5 mid-size model, Gemma 4 mid-size, Ministral 3 14B, gpt-oss-20b.
  • Evaluation: 40 English and 40 Arabic prompts rated by two bilingual editors on a 1–5 rubric, plus JSON validity on 30 brief prompts.
  • Result (illustrative): two candidates score within noise of each other; one produces invalid JSON 10% of the time and is dropped; the agency picks the smaller remaining model.

Hands-on: a mini evaluation harness

# eval_shortlist.py  (pip install openai)
import json, os, time
from openai import OpenAI

# Any OpenAI-compatible local server: Ollama, llama.cpp, LM Studio or vLLM
client = OpenAI(base_url=os.getenv("LOCAL_BASE_URL", "http://localhost:11434/v1"),
                api_key=os.getenv("LOCAL_API_KEY", "not-needed"))
MODELS = os.getenv("CANDIDATES", "qwen3:8b,gpt-oss:20b").split(",")
tasks = [json.loads(line) for line in open("tasks.jsonl", encoding="utf-8")]  # {"prompt":..., "must_include":[...]}

for model in MODELS:
    passed, latencies = 0, []
    for t in tasks:
        start = time.perf_counter()
        try:
            r = client.chat.completions.create(model=model, temperature=0,
                                               messages=[{"role": "user", "content": t["prompt"]}])
            text = r.choices[0].message.content or ""
        except Exception as e:
            print(f"{model}: error {e}")
            continue
        latencies.append(time.perf_counter() - start)
        passed += all(k.lower() in text.lower() for k in t["must_include"])
    if latencies:
        print(f"{model}: {passed}/{len(tasks)} passed, median latency {sorted(latencies)[len(latencies)//2]:.1f}s")

Keyword checks are crude; for real decisions add human rubric scores or an LLM judge (covered in the benchmarking lesson).

Pitfalls

  • Evaluating at full precision, then deploying a 4-bit quantization.
  • Picking by a public leaderboard that does not include your language or task.
  • Ignoring the chat template; the wrong template quietly degrades quality.

How to measure success

You can defend your choice with a table: candidates, constraints met, eval scores at deployment quantization, latency on target hardware, and the reason the smallest passing model won.

Video lecture: Model families in 2026 and how to shortlist

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

  1. Model families and how to shortlist
  2. Analogy: hiring for a role
  3. Size and architecture
  4. More dimensions
  5. Snapshot: September 2026
  6. Four-step shortlist
  7. Worked example: Manchester agency (illustrative)
  8. Hands-on harness
  9. Quiet failures
  10. Simple example: café feedback (illustrative)
  11. Plan for change
  12. Try this now
  13. Watch me do it
  14. Recap

Lecture transcript

Model families and how to shortlist

Every week a new open model claims the top spot somewhere. If you chase leaderboards you will never ship. So in this lesson you will learn the handful of dimensions that decide whether a model fits your job, get a snapshot of the major families as of September twenty twenty-six, and walk away with a four-step shortlist method and a tiny evaluation harness you can run today.

Analogy: hiring for a role

Choosing a model is a lot like hiring. You would not hire the candidate with the most impressive CV for every role. For a receptionist role you care about languages, manners and reliability. For an analyst role, reasoning. And you always run your own interview, because a CV is written to impress. Model cards and leaderboards are CVs. Your evaluation set is the interview. And just like hiring, the best candidate is often not the most senior one, it is the one who does this particular job well at a sensible cost.

Size and architecture

First, size. As a rough guide, one to four billion parameters suits phones and simple tasks, seven to fifteen billion suits laptops and many business tasks, twenty to forty billion gives workstation-class quality, and above a hundred billion is server territory. Second, dense versus mixture of experts. A dense model uses every parameter for every token. A mixture-of-experts model routes each token through a few experts. So it has total parameters, which must fit in memory, and active parameters, which drive speed. gpt-oss one-twenty B has about one hundred seventeen billion total but only about five billion active per token. Big memory, fast generation.

More dimensions

The other dimensions: reasoning mode, where thinking variants trade speed for better multi-step answers. Context length, where the advertised maximum is not the same as quality at that length. Modalities: text only, or images and audio too. Languages: if your users write Arabic, Urdu, or Roman Urdu mixed with English, test those specifically. Tool calling and structured output, which matter for agents. And license, which we covered last lesson.

Snapshot: September 2026

Now the snapshot, and please treat it as a snapshot. OpenAI's gpt-oss, twenty B and one-twenty B, are Apache licensed reasoning models shipped natively quantized. Google's Gemma 4 spans edge sizes to around thirty-one billion, now under Apache. Alibaba's Qwen three and three point five cover tiny to huge, mostly Apache. DeepSeek publishes very large mixture-of-experts reasoning models under MIT. Mistral 3 includes the Large 3 mixture-of-experts and the Ministral 3 edge models, all Apache. Meta's Llama 4 uses its community license. And Microsoft's Phi-4 family covers small models under MIT. Many other publishers ship strong open weights too.

Four-step shortlist

Here is the method. Step one, filter on hard constraints: license, languages, modalities, the memory you can afford, and the context you actually need. Step two, pick three to five candidates across sizes, one small, one mid, one large. Step three, run your own evaluation set, thirty to a hundred real tasks, at the exact quantization you will deploy. Step four, choose the smallest model that passes, then measure latency and cost on the real hardware.

Worked example: Manchester agency (illustrative)

Example. A Manchester agency drafts social posts in English and Arabic for Gulf clients and needs clean JSON campaign briefs. Constraints: permissive license, strong Arabic, a twenty-four gigabyte GPU per workstation. They test four candidates on forty English and forty Arabic prompts rated by bilingual editors, plus thirty JSON prompts. In this illustrative run, two models tie within noise, one breaks JSON too often and drops out, and they pick the smaller of the rest. That is the discipline in action.

Hands-on harness

The harness in your lesson is deliberately small. It talks to any OpenAI-compatible local server, which Ollama, llama dot cpp, LM Studio and vLLM all provide. It reads tasks from a JSON lines file, runs each candidate at temperature zero, checks for required keywords, and prints pass rate and median latency. Keyword checks are crude, so for real decisions you will add human ratings or a model judge, which we cover in the benchmarking lesson.

Quiet failures

Watch for three quiet failures. The first is evaluating a model at full precision on a borrowed cloud GPU, then deploying a four-bit version on a laptop and never re-testing. Quantization can change behavior, so evaluate what you ship. The second is trusting a public leaderboard that never tested your language or your task. The third is the chat template. Every model expects its conversation formatted in a specific way, and a runtime that applies the wrong template will quietly make a good model look mediocre. If one candidate looks strangely weak, check the template before you drop it.

Simple example: café feedback (illustrative)

A simple example before the agency one. A café chain in Karachi wants to classify customer feedback from its app into five buckets: taste, service, price, cleanliness and other. English and Roman Urdu. Constraints: runs on one small office PC with a sixteen gigabyte GPU. Candidates: a small Qwen model, Gemma 4's smaller variant, and Phi-4-mini. They label forty real comments, run the harness, and two models score almost the same. They pick the smaller one because it answers in under a second. Three candidates, forty examples, one afternoon.

Plan for change

Finally, plan for change. Whatever you choose today will be overtaken within months. So make switching cheap. Keep your prompts, evaluation set and harness independent of any one model. Use OpenAI-compatible endpoints so swapping is a configuration change. And schedule a quarterly re-shortlist: run the same evaluation on two or three new candidates and switch only when the evidence says so. That is how you benefit from the fast pace of open models without being whipped around by it.

Try this now

Try this now. Write down your hard constraints for one task: license, languages, modalities, memory budget and context length. Then list three candidate models of different sizes that meet them. Do not evaluate yet. Just the shortlist. You will be surprised how many models drop out at the constraint step, and how much time that saves before you run a single test.

Watch me do it

Watch me do it. My task: classify customer reviews for an e-commerce seller into five topics, in English and Arabic. Constraints first: permissive license, Arabic support, a sixteen gigabyte GPU. I pull three candidates of different sizes with Ollama, a small, a mid-size and a larger quantized model. I open my tasks file: forty reviews, each with the expected topic. I run the harness from the lesson against all three. The output prints pass rates and median latency. The small model gets thirty-one of forty, the mid-size thirty-seven, the larger thirty-eight but twice as slow. I read the three failures of the mid-size model: two are genuinely ambiguous reviews, one is a sarcastic Arabic comment. I decide on the mid-size model, add the sarcastic case to the tasks file for future comparisons, and write the result into a small table with license, pass rate and latency.

Recap

Recap. Pick models on dimensions, not headlines. Remember that mixture-of-experts models need memory for all their parameters but compute with a few. Know the major families, and verify every model card. Shortlist three to five, evaluate on your own tasks at deployment precision, and choose the smallest that passes. Your next step: write thirty real tasks from your job and run the harness on two local models.

Key takeaways

  • Choose on dimensions: size, dense vs MoE, reasoning, context, modality, languages, tool use, license
  • MoE models need memory for total parameters but compute scales with active parameters
  • Major 2026 families: gpt-oss, Gemma 4, Qwen3/3.5, DeepSeek, Mistral 3, Llama 4, Phi-4
  • Shortlist 3–5 candidates and test on your own tasks at deployment quantization
  • Pick the smallest model that passes

Try it

Write 30 real tasks from your work into tasks.jsonl, run the harness against two local models, and record pass rate and median latency in a shortlist table.