Emerging Tech Horizons: What's Next After Today's AIFrontier AI trends · Lesson 3 of 16
Frontier AI: reasoning models and agents
Video lecture
Frontier AI: reasoning models and agents
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Frontier AI
If you last looked closely at AI a couple of years ago, two things have changed dramatically. Models now reason, spending extra effort thinking before they answer. And they act, using tools and working through many steps towards a goal. In this lesson you'll understand both shifts, where they genuinely help, where they still fail, and what's worth watching next.
0:26 Reasoning models
Reasoning models first. OpenAI's o1, in September 2024, popularised training models to work through an extended chain of reasoning before giving a final answer. In January 2025, DeepSeek's R1 showed strong reasoning could be released with open weights, which intensified competition and pushed prices down. Today most major model families offer a reasoning or extended thinking mode, often with a dial for how much thinking to allow.
0:55 Why it matters
Why does this matter for your business? Because these two shifts change which tasks AI can take on. Two years ago, many teams tried AI on analysis and planning and were disappointed by shallow answers. Reasoning models changed that for many analytical tasks. And agents moved AI from drafting text to doing multi-step work in real systems. That means the list of things worth trying has grown, but so has the need for careful scoping, verification and access control. The organisations getting value are the ones that understand both the capability and its limits.
1:36 Corridor, notepad, laptop
Here's an analogy. A fast model is like asking a knowledgeable colleague a quick question in the corridor. A reasoning model is like asking them to sit down with a notepad for twenty minutes. For where's the meeting room, the corridor answer is perfect. For should we restructure our pricing, you want the notepad. And an agent is that colleague with a laptop, a phone and access to your systems, who can go and do the work, which is exactly why you give them clear instructions and check their output.
2:15 Simple example: two quotes
A simple example. You ask a fast model to compare two supplier quotes. It says supplier A is cheaper. You ask a reasoning model with the same quotes. It works through delivery charges, payment terms and a minimum order quantity buried in the small print, and concludes supplier B is cheaper for your actual volumes. Same question, same documents, different depth. That's test-time compute earning its cost, and it's why you check the reasoning, not just the conclusion.
2:49 Test-time compute
The idea to remember is test-time compute. You can buy a better answer on a hard problem by letting the model think longer, at the cost of time and money. That gives you a new design choice. Fast, cheap responses for routine work like rewriting an email. Slower, deeper reasoning for high-stakes analysis like comparing contracts or explaining a financial variance.
3:16 Strengths and blind spot
Reasoning shines on structured analysis, coding and debugging, planning, maths and logic, and critiquing its own drafts. But it has a blind spot. Reasoning can't fix wrong premises. Give a model an outdated price list and it will reason brilliantly to the wrong conclusion. It still needs current, correct information, through documents you provide or retrieval from trusted sources.
3:42 Agents
Now agents. An agent combines a model with tools, like search, code, APIs and browsers, with memory of what it has done, and with a loop: plan, act, observe, adjust. The Model Context Protocol made connecting models to tools and data much easier, and agent-to-agent protocols like A2A aim to let agents from different vendors work together. In short, AI has moved from answering questions to doing work.
4:12 Where agents deliver
Where are agents delivering real value today? In software engineering, coding agents read codebases, make multi-file changes, run tests and open pull requests for developers to review. In research, deep-research modes search, read and synthesise many sources. In operations, agents triage tickets, reconcile data and draft responses, with humans approving anything consequential. The pattern is consistent: well-scoped tasks, clear success criteria, good tools and human oversight.
4:41 Still risky
And where are they still risky? Open-ended autonomy on high-stakes tasks. Errors compound across many steps, and agents can be manipulated by content they read, a problem called prompt injection. That's why sensible teams scope tightly, restrict access, and keep humans approving actions that matter. Treat any claim of fully autonomous everything with healthy scepticism.
5:05 What to watch
What should you watch, as outlook rather than prediction? First, how long and complex a task agents can complete reliably, tracked by independent evaluations, not launch videos. If that keeps improving, more multi-hour knowledge work becomes delegable. Second, cost per unit of capability, which has historically fallen quickly; re-cost your shelved ideas every six months. Third, governance: the EU AI Act, sector regulators and standards like ISO 42001 will shape how agents can be deployed. Here's a real-world style example: a Dubai trading company uses reasoning mode for month-end variance analysis, but every explanation must cite ledger lines a human checks, and posting journals stays human-only.
5:51 When to use reasoning mode
A practical way to decide when to use reasoning mode is to ask two questions. Is the task hard, meaning multi-step, analytical or easy to get subtly wrong? And are the stakes high enough that a better answer is worth waiting and paying for? If both answers are yes, use reasoning. If either is no, a fast model is usually fine. Write that rule down for your team, and review it when new models arrive.
6:24 Three mistakes
Three common mistakes. First, using reasoning mode for everything, which is slow and expensive, or for nothing, which misses its value on hard problems. Second, trusting a long, confident chain of reasoning without checking its premises. Third, giving agents broad access before you've scoped the task and tested it. The safest path is narrow tasks, clear success criteria, least-privilege access, and humans approving anything consequential.
6:52 Try this now
Try this now. Pick three real tasks from your week: one routine, like rewriting an email; one analytical, like comparing two offers; and one planning task, like a ninety-day rollout. Run each in a fast mode and in a reasoning mode of the assistant you already use. Check the answers yourself for accuracy, note how long each took, and rate usefulness from one to five. Then write one sentence: we use reasoning mode when, and fast mode when. Pin that sentence where your team will see it, and revisit it when new models arrive.
7:33 Recap
To recap. Reasoning models trade time and cost for better answers on hard problems, but still need correct inputs. Agents add tools, memory and a loop, and work best when tightly scoped and overseen. Watch reliability, cost and governance using independent evidence. Your next step: run the reasoning-versus-fast test in the lesson text on three real tasks, and write one rule for when to use each mode. Next, multimodal AI, world models and on-device AI.
Two shifts that define today's frontier
Between 2024 and 2026, two developments changed what AI can do for a business. First, reasoning models: systems that spend extra computation "thinking" through a problem before answering, which markedly improved performance on maths, coding, analysis and multi-step planning. Second, agents: systems that use tools, take actions and work over many steps towards a goal, from coding agents that edit whole repositories to browser agents that operate websites. Understanding both, and their limits, is essential for any horizon scan.
Reasoning models: what changed
OpenAI's o1 (September 2024) popularised the approach of training models to produce an extended chain of reasoning before the final answer. DeepSeek's R1 (January 2025) showed that strong reasoning could be achieved and released with open weights, which intensified competition and lowered costs. Since then, most frontier model families (from Anthropic, Google, OpenAI and others) offer reasoning or "extended thinking" modes, often with controls for how much the model thinks.
The key idea for business leaders is test-time compute: you can buy better answers on hard problems by letting the model think longer, at the cost of latency and money. That creates a new design choice: fast, cheap responses for routine tasks; slower, deeper reasoning for high-stakes analysis.
Where reasoning models shine: structured analysis (financial models, contract comparisons), debugging and coding, planning multi-step work, maths and logic, and reviewing their own drafts.
Where they still fail: facts they were never given (they can reason confidently from wrong premises), tasks needing up-to-date information without retrieval, and judgement calls that depend on context the model lacks.
Agents: from answering to doing
An agent combines a model with tools (search, code execution, APIs, browsers, files), memory (what it has done and learned), and a loop (plan, act, observe, adjust). Standards such as the Model Context Protocol (MCP) made connecting models to tools and data much easier, and agent-to-agent protocols such as A2A aim to let agents from different vendors collaborate.
Where agents are delivering value as of 2026:
- Software engineering: coding agents that read codebases, make multi-file changes, run tests and open pull requests under developer review.
- Research and analysis: deep-research modes that search, read and synthesise many sources into a report.
- Operations: agents that triage tickets, reconcile data, draft responses, and operate software through APIs or screens, with human approval for consequential actions.
The consistent lesson from early deployments: agents work best on well-scoped tasks with clear success criteria, good tool access and human oversight. Open-ended autonomy on high-stakes tasks remains risky because errors compound and agents can be manipulated through the content they read (prompt injection).
Outlook: what to watch (reasoned, not predicted)
- Longer, more reliable autonomous work. Vendors and independent researchers track how long and complex a task an agent can complete reliably. If that trend continues, more multi-hour knowledge tasks become delegable. Watch independent evaluations rather than launch claims.
- Cost per unit of intelligence falling. Historically, the price of a given level of model capability has dropped quickly as newer, more efficient models arrive. If that continues, use cases that are uneconomic today may become viable. Re-cost your ideas every six months.
- Governance catching up. Rules such as the EU AI Act, sector regulators and standards (like ISO/IEC 42001) shape how agents can be deployed. Expect more requirements around transparency, oversight and logging.
Hands-on: a reasoning-versus-fast test
Pick three real tasks: one routine (rewrite an email), one analytical (compare two supplier quotes with different terms), one planning (draft a 90-day rollout). Run each on a fast model mode and a reasoning mode.
EVALUATION GRID
Task | Mode | Time taken | Cost (if visible) | Accuracy (checked by you) | Usefulness 1-5 | Notes
-----|------|------------|-------------------|---------------------------|----------------|------
Email rewrite | fast | ... | ... | ... | ... | ...
Email rewrite | reasoning | ...
Quote comparison | fast | ...
Quote comparison | reasoning | ...
90-day plan | fast | ...
90-day plan | reasoning | ...
Conclusion: use reasoning mode for ______ ; fast mode for ______Worked example: a Dubai finance team
A mid-size trading company in Dubai tested reasoning models on month-end variance analysis. The fast mode produced plausible but shallow explanations; the reasoning mode, given the ledger extracts, identified that a large variance came from a currency revaluation and a mis-coded invoice. The team adopted reasoning mode for analysis with a rule: every explanation must cite the specific ledger lines, which a human checks. They also built a small agent that pulls the extracts automatically, but kept posting journal entries as a human-only step.
Pitfalls
- Using reasoning mode for everything (slow and costly) or for nothing (missing its value).
- Trusting agent output without verification because the reasoning "looks thorough".
- Granting agents broad access before scoping and testing.
How to measure success
For each use case: accuracy on a checked sample, time saved, cost per task, and human review time. Re-run when vendors ship new models.
Key takeaways
- Reasoning models spend extra test-time compute to improve hard analysis, coding and planning, at a cost in latency and money.
- Agents combine models with tools, memory and a loop; MCP and A2A are making tool and agent connections standard.
- Agents deliver most value on well-scoped tasks with clear success criteria, good tools and human oversight.
- Watch reliability of long tasks, falling cost per capability and governance, using independent evidence rather than launch claims.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run the reasoning-versus-fast test on three real tasks and write one rule for when your team should use each mode.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.