Emerging Tech Horizons: What's Next After Today's AIPhysical constraints and digital trust · Lesson 10 of 16
Energy, compute and the physical limits of AI
Video lecture
Energy, compute and the physical limits of AI
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Energy, compute and AI
AI feels weightless. You type a question, an answer appears. But behind every answer is a data centre full of specialised chips drawing electricity, needing cooling, and connected to grids that are sometimes already stretched. In this lesson you'll learn why energy and compute are becoming strategic constraints, what the evidence says, and what that means for your costs, sustainability and choices.
0:27 The evidence (IEA, April 2025)
Let's look at the evidence. The International Energy Agency's Energy and AI report, published in April 2025, estimated that data centres used around four hundred and fifteen terawatt hours of electricity in 2024. In its base case, it projected that to more than double to around nine hundred and forty-five terawatt hours by 2030, just under three percent of global electricity, with AI the biggest driver. The IEA also stressed wide uncertainty. So treat that as a scenario, not a fact about the future.
1:04 Why it matters
Why does this matter to you? Because energy and compute constraints show up in three places you care about. Your costs: AI bills can grow quickly as usage spreads across teams. Your commitments: if you report emissions or have sustainability targets, AI use now contributes to them. And your options: data residency rules and regional capacity decide which models and providers you can actually use. Leaders who understand the physical side of AI make better choices about which workloads to run, where, and on which models.
1:41 Like air travel
Here's an analogy. AI is a bit like air travel. The ticket price per journey has fallen a lot over the decades, thanks to more efficient planes. But because flying got cheaper, far more people fly, and total fuel use grew. AI is showing a similar pattern: each answer tends to get cheaper and more efficient, while total demand for compute and electricity keeps rising. Both things are true at once, and your planning needs to hold both.
2:15 Simple example: sorting emails
A simple example of right-sizing. A customer-service team sends every incoming email to a large reasoning model just to label it as billing, delivery or complaint. That's like hiring a senior lawyer to sort the post. Routing those labels to a small, fast model costs a fraction as much and is usually just as accurate for simple categories. The large model is then reserved for the tricky complaints that genuinely need careful reasoning.
2:47 Other constraints
Electricity isn't the only constraint. Advanced chips depend on a concentrated supply chain, leading-edge fabrication, advanced packaging and high-bandwidth memory, which is exposed to export controls and geopolitics. In many regions, new data centres wait a long time for grid connections. Hot climates, including the Gulf, raise cooling demands, pushing operators towards liquid cooling and water-efficient designs. And communities scrutinise land, water and local power prices.
3:16 Efficiency vs demand
Now the tension that confuses many forecasts. On one side, efficiency keeps improving: better chips, better model designs, smaller distilled models, quantisation, caching and smarter routing all cut the energy and cost per unit of capability. On the other side, cheaper AI unlocks more use, and reasoning models and agents consume much more compute per task. Economists call this the rebound effect, or Jevons paradox. Which force wins can change from year to year.
3:48 Outlook
What's reasonably robust, as outlook? Energy and compute will stay strategic constraints for AI providers. They'll shape where data centres are built, in regions with available power, including the Gulf states, which are investing heavily in AI infrastructure. They'll shape pricing. And they'll shape sovereignty debates, as governments in Saudi Arabia, the UAE, the EU and elsewhere pursue in-region AI capacity.
4:15 What it means for you
What does this mean for you? Cost first: unit prices can fall while your total bill rises with usage, so budget for growth and watch cost per task. Sustainability: if you report emissions, AI use flows into Scope 3 through your cloud providers, so ask them for data. Resilience: data residency rules and capacity limits may restrict which regions and providers you can use. And design: right-size models, cache repeated work, batch non-urgent jobs.
4:47 Three levers
Here's a quick hands-on. List your AI workflows with the model used, calls per month, tokens, cost, whether a smaller model could do it, and whether it's cacheable or batchable. Then apply three levers. Route classification, extraction and short rewrites to a small model. Cache long, repeated prompts like brand guidelines. And use batch APIs, often discounted, for non-urgent jobs like nightly summaries. The lesson text has a review template and a simple router sketch.
5:20 Worked example: UK retailer
A UK retailer saw its AI bill triple in a year as product descriptions, service summaries and search features grew. The review showed most calls were simple extractions going to a large model. Routing them to a smaller model, caching the long brand-guidelines prompt, and batching nightly catalogue jobs cut cost per task substantially, with no measurable quality drop in spot checks. They also added AI usage to their Scope 3 data requests to cloud providers.
5:53 Region and residency
Where your AI runs also matters for compliance. Data residency rules in Saudi Arabia, the UAE and the EU can restrict where personal data is processed, and some sectors have stricter rules. Many providers now offer in-region processing. When you pick a provider, check the regions available for the models you need, because the newest models are not always offered in every region at launch.
6:21 Three mistakes
Three common mistakes. First, assuming falling unit prices mean falling total costs. Usage usually grows. Second, using the biggest reasoning model for trivial tasks. Third, choosing providers and regions without checking data residency rules and capacity, which can force painful migrations later, especially for teams operating across the EU, the UK and the Gulf.
6:44 Try this now
Try this now. List your five busiest AI workflows. For each, note the model or tier used, roughly how many calls per month, whether a smaller model could do the job, whether the prompts repeat and could be cached, and whether the work could run overnight in a batch. Pick the single workflow with the most obvious saving, apply one change, and measure cost per task and quality for two weeks. Then ask your main AI provider what emissions and region information they can share for your usage.
7:22 Recap
To recap. AI runs on chips, power, cooling and grids, and those limits shape its cost and geography. The IEA's base case sees data-centre electricity more than doubling by 2030, but treat that as a scenario. Efficiency and demand pull against each other. For your organisation: watch cost per task, right-size, cache and batch, and include AI in sustainability and residency decisions. Your next step: complete the workload review. Next: synthetic media and provenance.
AI runs on power, chips and water
It is easy to think of AI as weightless software. In reality, every model is trained and served in data centres packed with specialised chips (GPUs and other accelerators), which need large amounts of electricity, cooling (often water-intensive), network capacity and land. These physical constraints increasingly shape AI's pace, cost and geography, and they belong in any serious horizon scan.
What the evidence says
The International Energy Agency's Energy and AI report (April 2025) estimated that data centres used around 415 TWh of electricity in 2024 and, in its base case, projected this to more than double to around 945 TWh by 2030, just under 3% of global electricity consumption, with AI the most important driver. The United States and China account for most of the projected growth. The IEA also noted wide uncertainty ranges, depending on efficiency gains, chip supply and adoption. Treat any single projection as a scenario, not a fact about the future.
Beyond electricity:
- Chips: advanced accelerators depend on a concentrated supply chain (leading-edge fabrication, advanced packaging, high-bandwidth memory), subject to export controls and geopolitical risk.
- Grid connections: in many regions, the wait for new grid capacity is a bottleneck for new data centres.
- Cooling and water: hot climates (including the Gulf) raise cooling demands; operators increasingly use liquid cooling and water-efficient designs.
- Local impact: communities and regulators scrutinise land use, water, noise and electricity prices.
Why efficiency keeps surprising people
Two forces pull in opposite directions:
- Efficiency gains: newer chips, better model architectures, smaller distilled models, quantisation, caching and smarter routing mean the cost and energy per unit of capability tend to fall.
- Demand growth: cheaper AI unlocks more use (reasoning models and agents also use much more compute per task), which can raise total consumption. Economists call this the rebound effect or Jevons paradox.
The outlook depends on which force dominates in each period. What is fairly robust: energy and compute will be strategic constraints for AI providers, shaping where data centres are built (regions with available power, including the Gulf states, which are investing heavily in AI infrastructure) and how models are priced.
What this means for your organisation
- Cost: model prices can fall per unit, but your total bill can rise with usage. Budget for growth, monitor per-task cost, and route simple tasks to smaller models.
- Sustainability reporting: if you report emissions (for example under the EU's CSRD for in-scope companies, or voluntary frameworks), AI use contributes to Scope 3 emissions via cloud providers. Ask providers for emissions data and region-level information.
- Resilience and sovereignty: data-residency rules and capacity constraints may affect which regions and providers you can use. Governments in KSA and the UAE, the EU and elsewhere are pursuing sovereign or in-region AI capacity.
- Design choices: right-size models, cache repeated work, batch non-urgent jobs, and avoid "reasoning on everything".
Hands-on: an AI cost-and-footprint review
AI WORKLOAD REVIEW (one row per workflow)
Workflow | Model/tier | Calls per month | Avg tokens in/out | Cost/month | Could a smaller model do it? | Cacheable? | Batchable? | Provider region | Emissions data available?Then apply three optimisations and re-measure:
- Routing: send classification, extraction and short rewrites to a small, fast model; reserve large or reasoning models for hard tasks.
- Caching: use provider prompt caching for long, repeated system prompts and documents.
- Batching: use providers' batch APIs (often discounted) for non-urgent jobs such as nightly summaries.
# Simple router sketch (pseudo-configurable): pick a model tier by task type
ROUTES = {"classify": "small", "extract": "small", "rewrite_short": "small",
"analyse": "large", "plan": "reasoning"}
def pick_model(task_type: str) -> str:
tier = ROUTES.get(task_type, "large")
return {"small": "SMALL_MODEL_ID", "large": "LARGE_MODEL_ID",
"reasoning": "REASONING_MODEL_ID"}[tier] # set IDs from your provider's current docsWorked example: a UK e-commerce retailer
A UK retailer's AI bill tripled in a year as product-description generation, customer-service summaries and search features grew. A workload review showed most calls were simple extractions sent to a large model. Routing them to a smaller model, caching the long brand-guidelines prompt and batching nightly catalogue jobs cut cost per task substantially without a measurable quality drop in their spot checks. The retailer also added AI usage to its Scope 3 data requests to cloud providers.
Pitfalls
- Assuming falling unit prices mean falling total costs.
- Using the largest reasoning model for trivial tasks.
- Ignoring data-residency and capacity constraints in provider choice.
How to measure success
Cost per task, share of calls routed to smaller models, cache hit rate, and provider-reported emissions or energy data where available.
Key takeaways
- AI depends on chips, electricity, cooling, grid connections and land; these physical limits shape pace, cost and location.
- The IEA's base case projects data-centre electricity more than doubling from about 415 TWh (2024) to about 945 TWh (2030); treat projections as scenarios.
- Efficiency gains and demand growth pull in opposite directions; total AI costs can rise even as unit prices fall.
- Route tasks to right-sized models, cache and batch, and include AI in sustainability and residency decisions.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Complete the AI workload review for your top five AI workflows and apply routing, caching or batching to at least one, then re-measure cost per task.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.