Emerging Tech Horizons: What's Next After Today's AIFrontier AI trends · Lesson 3 of 16

Frontier AI: reasoning models and agents

Article · 7 min · 8 min lecture

Video lecture

Frontier AI: reasoning models and agents

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Frontier AI

  • Reasoning models
  • Agents that act
  • Where they help and fail
  • What to watch

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Two shifts that define today's frontier

Between 2024 and 2026, two developments changed what AI can do for a business. First, reasoning models: systems that spend extra computation "thinking" through a problem before answering, which markedly improved performance on maths, coding, analysis and multi-step planning. Second, agents: systems that use tools, take actions and work over many steps towards a goal, from coding agents that edit whole repositories to browser agents that operate websites. Understanding both, and their limits, is essential for any horizon scan.

Reasoning models: what changed

OpenAI's o1 (September 2024) popularised the approach of training models to produce an extended chain of reasoning before the final answer. DeepSeek's R1 (January 2025) showed that strong reasoning could be achieved and released with open weights, which intensified competition and lowered costs. Since then, most frontier model families (from Anthropic, Google, OpenAI and others) offer reasoning or "extended thinking" modes, often with controls for how much the model thinks.

The key idea for business leaders is test-time compute: you can buy better answers on hard problems by letting the model think longer, at the cost of latency and money. That creates a new design choice: fast, cheap responses for routine tasks; slower, deeper reasoning for high-stakes analysis.

Where reasoning models shine: structured analysis (financial models, contract comparisons), debugging and coding, planning multi-step work, maths and logic, and reviewing their own drafts.

Where they still fail: facts they were never given (they can reason confidently from wrong premises), tasks needing up-to-date information without retrieval, and judgement calls that depend on context the model lacks.

Agents: from answering to doing

An agent combines a model with tools (search, code execution, APIs, browsers, files), memory (what it has done and learned), and a loop (plan, act, observe, adjust). Standards such as the Model Context Protocol (MCP) made connecting models to tools and data much easier, and agent-to-agent protocols such as A2A aim to let agents from different vendors collaborate.

Where agents are delivering value as of 2026:

  • Software engineering: coding agents that read codebases, make multi-file changes, run tests and open pull requests under developer review.
  • Research and analysis: deep-research modes that search, read and synthesise many sources into a report.
  • Operations: agents that triage tickets, reconcile data, draft responses, and operate software through APIs or screens, with human approval for consequential actions.

The consistent lesson from early deployments: agents work best on well-scoped tasks with clear success criteria, good tool access and human oversight. Open-ended autonomy on high-stakes tasks remains risky because errors compound and agents can be manipulated through the content they read (prompt injection).

Outlook: what to watch (reasoned, not predicted)

  • Longer, more reliable autonomous work. Vendors and independent researchers track how long and complex a task an agent can complete reliably. If that trend continues, more multi-hour knowledge tasks become delegable. Watch independent evaluations rather than launch claims.
  • Cost per unit of intelligence falling. Historically, the price of a given level of model capability has dropped quickly as newer, more efficient models arrive. If that continues, use cases that are uneconomic today may become viable. Re-cost your ideas every six months.
  • Governance catching up. Rules such as the EU AI Act, sector regulators and standards (like ISO/IEC 42001) shape how agents can be deployed. Expect more requirements around transparency, oversight and logging.

Hands-on: a reasoning-versus-fast test

Pick three real tasks: one routine (rewrite an email), one analytical (compare two supplier quotes with different terms), one planning (draft a 90-day rollout). Run each on a fast model mode and a reasoning mode.

EVALUATION GRID
Task | Mode | Time taken | Cost (if visible) | Accuracy (checked by you) | Usefulness 1-5 | Notes
-----|------|------------|-------------------|---------------------------|----------------|------
Email rewrite | fast | ... | ... | ... | ... | ...
Email rewrite | reasoning | ...
Quote comparison | fast | ...
Quote comparison | reasoning | ...
90-day plan | fast | ...
90-day plan | reasoning | ...
Conclusion: use reasoning mode for ______ ; fast mode for ______

Worked example: a Dubai finance team

A mid-size trading company in Dubai tested reasoning models on month-end variance analysis. The fast mode produced plausible but shallow explanations; the reasoning mode, given the ledger extracts, identified that a large variance came from a currency revaluation and a mis-coded invoice. The team adopted reasoning mode for analysis with a rule: every explanation must cite the specific ledger lines, which a human checks. They also built a small agent that pulls the extracts automatically, but kept posting journal entries as a human-only step.

Pitfalls

  • Using reasoning mode for everything (slow and costly) or for nothing (missing its value).
  • Trusting agent output without verification because the reasoning "looks thorough".
  • Granting agents broad access before scoping and testing.

How to measure success

For each use case: accuracy on a checked sample, time saved, cost per task, and human review time. Re-run when vendors ship new models.

Key takeaways

  • Reasoning models spend extra test-time compute to improve hard analysis, coding and planning, at a cost in latency and money.
  • Agents combine models with tools, memory and a loop; MCP and A2A are making tool and agent connections standard.
  • Agents deliver most value on well-scoped tasks with clear success criteria, good tools and human oversight.
  • Watch reliability of long tasks, falling cost per capability and governance, using independent evidence rather than launch claims.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What does 'test-time compute' mean for a business user?
  2. Where do agents currently deliver the most reliable value?
  3. A reasoning model gives a detailed, confident analysis based on an outdated price list. What went wrong?

Put it into practice

Run the reasoning-versus-fast test on three real tasks and write one rule for when your team should use each mode.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.