Skip to content

Open-Weight and Local AI: Run, Choose and Deploy Your Own Models · Quantization and hardware sizing · lesson 6 of 16 · 14 min

Hardware choices: laptops, workstations, servers and cloud GPUs

Start from the workload, not the chip

Hardware decisions are sizing decisions (previous lesson) plus three practical constraints: budget, where the data must live, and who will operate it. Specs and prices change fast, so this lesson teaches categories and trade-offs; check current vendor specifications and local availability (import lead times in Pakistan, the UAE and Saudi Arabia can differ a lot from the US or UK).

The five tiers

| Tier | Typical memory for models | Good for | Watch out for | |---|---|---|---| | CPU-only server or PC | System RAM (lots, cheap) | Small models, batch jobs, embeddings, low traffic | Slow generation; fine for background work | | Apple silicon Mac | Unified memory shared by CPU and GPU (a portion usable for models) | Personal assistants, prototyping, running mid-size models quietly; MLX and llama.cpp run well | Throughput for many users is limited; not a multi-user server | | Consumer NVIDIA GPU (e.g. 16–32 GB cards) | Dedicated VRAM | Single-developer inference, 7–32B models at 4-bit, fine-tuning small models with LoRA | VRAM ceiling; multi-GPU consumer setups add complexity | | Workstation GPU (e.g. 48–96 GB class) | Large VRAM, ECC, pro drivers | Team servers for 10–50 users, bigger models, on-prem compliance | Cost, power, cooling | | Data-center GPU / cloud (e.g. 80 GB+ accelerators from NVIDIA or AMD) | High-bandwidth memory, fast interconnects | High-concurrency serving, very large MoE models, FP8 | Cost; availability; cloud region choice for data residency |

Key trade-offs explained

Memory capacity vs bandwidth. Capacity decides what fits; bandwidth decides how fast it generates. A machine with huge but slower memory can load a big MoE model yet generate more slowly than a smaller, faster GPU running a smaller model.

Unified memory (Apple silicon, some AMD APUs). The GPU can address a large share of system memory, so a high-memory Mac can load models that would need multiple consumer GPUs. It is excellent for individuals and small offices; for many concurrent users, dedicated GPUs with batching servers win.

MoE changes the maths. Very large MoE models need huge memory (all experts) but relatively little compute per token. That makes high-memory machines and CPU+GPU offloading (keeping some experts in system RAM) attractive for low-concurrency use.

Multi-GPU. Tensor parallelism splits each layer across GPUs (fast interconnect matters); pipeline parallelism splits layers across GPUs. vLLM and SGLang support both. For a single team server, one larger card is operationally simpler than several small ones.

Power and cooling. A high-end GPU workstation can draw substantial power under load. In hot climates (Karachi, Dubai, Riyadh summers) plan cooling and power backup; thermal throttling silently cuts throughput.

Cloud vs on-prem. Cloud GPUs avoid capital cost and let you scale up for a launch, then down. On-prem wins on steady utilization, strict data locality and predictable cost. Many organizations prototype in the cloud (in a region that meets residency needs, for example a UAE, Saudi or UK region where the provider offers GPUs) and move steady workloads on-prem later.

Worked example: three buyers, three answers

  1. Solo consultant in London. Needs a private writing and document assistant. A high-memory Apple silicon laptop runs a 14–32B model at 4-bit quietly on battery. No server needed.
  2. 30-person agency in Lahore. Needs a shared assistant for drafting and internal RAG, about 15 concurrent users at peak. One workstation with a 48 GB-class GPU running vLLM and a 14–32B model at 4-bit, plus a UPS. Cloud burst for rare peaks.
  3. Bank in Riyadh. Regulated data, hundreds of concurrent users, strict on-prem policy. A small on-prem cluster of data-center GPUs serving FP8 models with vLLM or SGLang, behind the bank's gateway, with a formal capacity plan.

Hands-on: a total cost of ownership sketch

Use this template to compare options over 24 months. All figures are placeholders to replace with your quotes.

Option: Workstation (48 GB-class GPU)          Option: Cloud GPU (on-demand)
Hardware (one-off)            = [quote]        Hourly rate × hours used      = [quote × hours]
Power: watts × hours × tariff = [calc]         Storage + egress              = [quote]
Cooling / UPS                 = [quote]        Reserved/committed discount   = [if any]
Admin time: hours/month × rate × 24 = [calc]   Admin time (less, not zero)   = [calc]
Depreciation / resale         = [estimate]
-----------------------------------------------------------------------------
24-month total                = ...            24-month total                = ...
Cost per 1,000 requests at expected volume     Cost per 1,000 requests at expected volume

Then compare both with an API model's cost per 1,000 requests at the same quality. If the API is cheaper and data rules allow it, use the API until volume grows.

Pitfalls

  • Buying on peak FLOPS numbers when your bottleneck is memory capacity or bandwidth.
  • Ignoring utilization: a GPU that is busy 5% of the time is expensive per request.
  • Forgetting cooling, power backup and noise in small offices.
  • Choosing a cloud region for price when your data must stay in-country.

How to measure success

You can show a sizing estimate, a 24-month cost comparison against at least one alternative (cloud or API), and a utilization forecast that justifies the purchase.

Video lecture: Hardware choices: laptops, workstations, servers and cloud GPUs

Lecture coming soon · 14 chapters · about 9 minutes. Read the full transcript below.

  1. Hardware choices
  2. Analogy: delivery vehicles
  3. Five tiers
  4. Capacity vs bandwidth
  5. MoE and multi-GPU
  6. The unglamorous parts
  7. Three buyers
  8. Hands-on: 24-month TCO
  9. Simple example: two-person studio
  10. Rent before you buy
  11. FAQ
  12. Try this now
  13. Watch me do it
  14. Recap

Lecture transcript

Hardware choices

Two companies buy the same expensive GPU. One runs it flat out serving hundreds of staff and pays it back in months. The other runs it five percent of the time and quietly burns money for two years. Same chip, opposite outcomes. In this lesson you will learn the five hardware tiers for running models, the trade-offs that actually matter, and a simple total cost sketch so your purchase is defensible.

Analogy: delivery vehicles

Buying AI hardware is like choosing vehicles for a delivery business. A motorbike is cheap and nimble, perfect for one parcel at a time: that is a laptop. A van carries a team's worth of parcels: that is a workstation GPU. A truck fleet moves a whole city's deliveries: that is a data-center cluster. Buying a truck to deliver one parcel a day wastes money, and sending a motorbike to move a warehouse fails. Match the vehicle to the load, and remember fuel, insurance and parking, which here are power, cooling and people.

Five tiers

Tier one is CPU only: slow generation, but fine for small models, embeddings and background batch work. Tier two is Apple silicon, where the GPU shares a large pool of unified memory, great for individuals and prototypes. Tier three is a consumer NVIDIA card with sixteen to thirty-two gigabytes, perfect for one developer running models up to about thirty billion at four bits. Tier four is a workstation GPU with around forty-eight to ninety-six gigabytes, the sweet spot for a team server. And tier five is data-center accelerators and cloud GPUs for high concurrency and giant models. Specs change fast, so check current vendor sheets.

Capacity vs bandwidth

The first trade-off is capacity versus bandwidth. Capacity decides whether a model fits. Bandwidth decides how fast it talks. A machine with a huge but slower memory pool can load a big model, yet generate more slowly than a smaller, faster GPU running a smaller model. So always ask two questions: will it fit, and how fast will it read.

MoE and multi-GPU

Mixture-of-experts models bend the rules. They need memory for every expert, but only a few experts work on each token. That makes high-memory machines, and setups that keep some experts in regular system memory, surprisingly usable for one or a few users. For many concurrent users, though, dedicated GPUs with a batching server still win. Multi-GPU setups can split a model across cards, but for one team server, one bigger card is usually simpler to run than several small ones.

The unglamorous parts

Now the unglamorous parts that decide real outcomes. Power and cooling: a high-end workstation under load draws serious power, and in a Karachi or Riyadh summer, poor cooling means silent throttling. Plan a UPS and airflow. Cloud versus on-premise: cloud avoids upfront cost and scales for launches, on-premise wins with steady use, strict data locality and predictable bills. If data must stay in-country, choose a cloud region that meets that requirement, not just the cheapest one.

Three buyers

Three buyers, three answers. A solo consultant in London needs a private writing assistant: a high-memory Apple silicon laptop runs a fourteen to thirty-two billion model quietly on battery. A thirty-person agency in Lahore needs a shared assistant for about fifteen concurrent users: one workstation GPU, a batching server, and a UPS, with occasional cloud burst. A bank in Riyadh with hundreds of users and a strict on-premise policy needs a small cluster of data-center GPUs behind its own gateway, with a formal capacity plan.

Hands-on: 24-month TCO

Your hands-on tool is a twenty-four month cost sketch. On one side, hardware, power at your local tariff, cooling and UPS, admin hours, and resale value. On the other, cloud hourly rates times hours, storage, any committed-use discount, and admin time. Divide each by expected volume to get cost per thousand requests, and compare it with an API model at the same quality. If the API is cheaper and your data rules allow it, keep using the API until volume grows.

Simple example: two-person studio

A simple example. A two-person design studio in Dubai wants a private assistant for writing proposals. Two users, rarely at the same time, a fourteen billion model at four bits is plenty. Their options: a high-memory Apple silicon laptop they would buy anyway, or a cloud GPU rented by the hour. The laptop wins: it already exists in their budget, runs quietly, and keeps client proposals in the office. A GPU server would sit idle most of the day. Sometimes the right hardware decision is to buy nothing new at all.

Rent before you buy

One more practical tip before you buy anything: rent first. Most cloud providers let you rent the exact GPU class you are considering by the hour. Spend a day running your real prompts at your expected concurrency, with the same engine and quantization you plan to use, and record throughput, latency and memory. That one day of rental often costs less than a single wrong component, and it turns your purchase request from a guess into a measured plan. If data rules prevent you from using real prompts in the cloud, use synthetic prompts with the same length and structure.

FAQ

A question I get all the time: should we buy one big GPU or two smaller ones? For most single-team servers, one bigger card is simpler. The whole model fits in one place, you avoid splitting layers across cards, and there is one thing to monitor and cool. Two smaller cards make sense when the model does not fit on any single card you can buy, or when you want two independent replicas for resilience. And a second common question: is a Mac good enough for our team? For a few people using it one at a time, often yes. For twenty people hitting it at once, a dedicated GPU with a batching server will serve them far better.

Try this now

Try this now. Write down your expected number of concurrent users at peak, and the context length they really need. Then place yourself on the five tiers. If you land on tier two or three, you probably do not need a server yet. If you land on tier four or five, your next step is to rent that exact GPU class in the cloud for a day and measure, before you buy anything.

Watch me do it

Watch me do it. A thirty-person office asks me what to buy. I start with numbers, not shopping. Peak concurrent users: I check their chat tool logs and see about twelve at the busiest hour. Context: most documents fit in eight thousand tokens. I run the sizing calculator for a fourteen billion model at four bits: comfortably inside a forty-eight gigabyte workstation card, with room for more users. Next I rent that GPU class in the cloud for an afternoon, serve the model with vLLM, and replay fifty real prompts at twelve concurrent requests. Latency looks good. Then the cost sketch: workstation price, power at their local tariff, a UPS, and two hours a month of admin, over twenty-four months, divided by expected requests. I compare with their current API bill. The workstation pays for itself in the second year, and it keeps client files in the office. That is the recommendation I send.

Recap

Recap. Start from sizing and the workload. Capacity decides fit, bandwidth decides speed. Unified memory is great for individuals, batching GPU servers are best for teams. Mixture-of-experts models need memory more than compute. And judge purchases on twenty-four month cost per thousand requests, including power, cooling and people. Your next step: fill in the cost template for two options and compare them with an API you already use.

Key takeaways

  • Capacity decides what fits; bandwidth decides how fast it generates
  • Apple silicon unified memory is great for individuals; batching servers on dedicated GPUs win for teams
  • MoE models need big memory but little compute per token
  • Compare on 24-month cost per 1,000 requests, including power, cooling and admin time
  • Choose cloud regions for data residency, not just price

Try it

Fill in the 24-month TCO template for two options (on-prem vs cloud) and compare the cost per 1,000 requests with an API model you already use.