Latest AI Techniques: RAG, Tool Use, Agents & MCPTool and function calling · Lesson 6 of 20

Function calling fundamentals

Article · 12 min · 9 min lecture

Video lecture

Function calling fundamentals

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Tool use (function calling)

  • The model requests; your code executes
  • Defining tools well
  • Your code is the security boundary

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

What tool use is

Tool use (also called function calling) lets a model request that your application run a function: search a database, call an API, do a calculation, send a message. The model does not execute anything itself. It outputs a structured request (the tool name and arguments), your code executes it, and you return the result to the model, which then continues.

This turns a model from a text generator into the reasoning layer of a system that can fetch live data and take actions.

The request cycle

1. You send: messages + tool definitions (name, description, input schema)
2. Model replies: either a normal answer, or a tool call such as
   get_order_status({"order_id": "A1234"})
3. Your code: validates arguments, checks permissions, executes the function
4. You send back: the tool result, linked to that call
5. Model replies: a final answer using the result, or another tool call

Many APIs allow several tool calls in one turn (for example fetching three records in parallel), and let you control whether the model must use a tool, may use one, or must use a specific one.

Defining a tool

A tool definition is a name, a natural-language description and a JSON Schema for its inputs:

{
  "name": "get_order_status",
  "description": "Look up the current status and estimated delivery date of a customer's order. Use when the customer asks where their order is or when it will arrive. Requires the order ID, which starts with 'A' followed by 4-8 digits.",
  "input_schema": {
    "type": "object",
    "properties": {
      "order_id": {"type": "string", "pattern": "^A[0-9]{4,8}$",
                   "description": "Order ID from the confirmation email, e.g. A1234"}
    },
    "required": ["order_id"]
  }
}

The exact field names differ slightly between providers, but the ingredients are the same everywhere.

Descriptions are prompts

The model decides whether and how to call a tool mainly from its name and description. Good descriptions state:

  • what the tool does and returns
  • when to use it (and when not to)
  • what inputs mean, with formats and examples
  • limitations ("returns at most 20 results", "data refreshed hourly")

Vague descriptions ("Gets data") cause missed calls, wrong calls and invented arguments.

Executing safely

Your code is the security boundary. The model's tool call is a request, and it may be wrong or manipulated (for example via prompt injection in earlier content). So:

  • Validate arguments against the schema and business rules.
  • Authorise using the real user's permissions, not the model's say-so. If a user can only see their own orders, enforce that in code.
  • Confirm consequential actions with the user (refunds, sending messages, deleting data).
  • Rate-limit and time-out external calls.
  • Log every call and result.

Returning results well

  • Return concise, structured results. A 5,000-line API response wastes context and distracts the model; extract what matters.
  • Return useful errors: "Order A1234 not found. Ask the customer to check the ID in their confirmation email" helps the model recover. A raw stack trace does not.
  • Mark untrusted content (for example text from web pages or emails) so the model treats it as data.

Worked example: an order-tracking assistant

User: Where's my order? It's A20931.
Model -> get_order_status({"order_id": "A20931"})
App   -> checks that A20931 belongs to the logged-in customer; calls the
         order system; returns {"status": "in_transit", "carrier": "...",
         "eta": "2026-09-29"}
Model -> "Your order is on its way and is expected on 29 September..."

If the customer asks about someone else's order, the app's permission check fails and returns "Not authorised for this order", and the model explains politely. The model never had the power to bypass that check.

Failure modes

  • Hallucinated arguments: the model invents an order ID when the user didn't provide one. Mitigate with schema patterns, descriptions ("ask the user if missing") and validation.
  • Tool overload: dozens of similar tools confuse selection. Consolidate, or load tools dynamically based on the task.
  • Silent failures: a tool returns empty results and the model answers as if it had data. Return explicit "no results" messages.

Hands-on: a complete tool-use loop with validation and authorisation

Below is a manual tool loop with the Claude API (Python SDK). It shows every step you read about: sending a tool definition, receiving a tool_use block, validating and authorising in your own code, returning a tool_result (with is_error when something fails) and letting the model finish. The SDK also offers a beta "tool runner" helper that drives this loop for you; writing it once by hand is the best way to understand what such helpers do.

import json, os, re
import anthropic

client = anthropic.Anthropic()                       # reads ANTHROPIC_API_KEY
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")  # check current models in the docs

TOOLS = [{
    "name": "get_order_status",
    "description": ("Look up status and estimated delivery of ONE order belonging to the signed-in "
                    "customer. Use when they ask where an order is. If no order ID was given, ask "
                    "for it instead of guessing. IDs look like A1234 (A + 4-8 digits)."),
    "input_schema": {"type": "object",
                     "properties": {"order_id": {"type": "string", "pattern": "^A[0-9]{4,8}$"}},
                     "required": ["order_id"], "additionalProperties": False},
}]

FAKE_DB = {"A20931": {"owner": "cust_17", "status": "in_transit", "eta": "2026-09-29"}}

def get_order_status(order_id: str, customer_id: str) -> dict:
    if not re.fullmatch(r"A[0-9]{4,8}", order_id):
        raise ValueError("order_id must look like A1234; ask the customer to check their email")
    order = FAKE_DB.get(order_id)
    if order is None or order["owner"] != customer_id:       # authorise in code, not in the prompt
        raise PermissionError("No order with that ID for this customer")
    return {"order_id": order_id, "status": order["status"], "eta": order["eta"]}

def run(user_text: str, customer_id: str, max_turns: int = 6) -> str:
    messages = [{"role": "user", "content": user_text}]
    for _ in range(max_turns):
        resp = client.messages.create(model=MODEL, max_tokens=4000, tools=TOOLS,
                                      system="You are a concise order-support assistant.",
                                      messages=messages)
        messages.append({"role": "assistant", "content": resp.content})
        if resp.stop_reason != "tool_use":
            return "".join(b.text for b in resp.content if b.type == "text")
        results = []
        for block in resp.content:
            if block.type != "tool_use":
                continue
            try:
                out = get_order_status(block.input["order_id"], customer_id)
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": json.dumps(out)})
            except (ValueError, PermissionError, KeyError) as err:
                results.append({"type": "tool_result", "tool_use_id": block.id,
                                "content": str(err), "is_error": True})
        messages.append({"role": "user", "content": results})  # all results in ONE message
    return "Sorry, I could not complete that. A colleague will follow up."

print(run("Where's my order A20931?", customer_id="cust_17"))
print(run("Where's order A20931?", customer_id="cust_99"))   # authorisation failure path

Notice three production habits: the permission check uses the real signed-in customer, not anything the model says; errors are returned as readable tool_result content with is_error so the model can recover gracefully; and the loop has a turn limit.

The same idea on other platforms

OpenAI's Responses API, Google's Gemini API and open-source runtimes all implement the same cycle with slightly different field names (for example function_call and function_call_output items rather than tool_use and tool_result blocks). Frameworks such as the OpenAI Agents SDK wrap it in helpers. Learn the cycle once and every SDK becomes readable. The flagship course Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API covers provider differences, streaming tool calls and production error handling in depth.

Going further

Some providers offer server-side tools (web search, web fetch, code execution) that run in the provider's environment, alongside client-side tools you execute. The design principles are the same: clear descriptions, least privilege, and treating every result as input to be checked.

Key takeaways

  • The model requests tool calls; your code validates, authorises, executes and returns results.
  • Tool names, descriptions and schemas are prompts: state what, when, input formats and limits.
  • Your code is the security boundary: validate arguments, use the real user's permissions, confirm consequential actions.
  • Return concise, structured results and helpful errors; flag untrusted content.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Who actually executes a function when a model makes a tool call?
  2. A model calls get_order_status for an order that belongs to another customer. What should prevent data exposure?
  3. A tool returns a 5,000-line JSON payload. What is the better practice?

Put it into practice

Write a tool definition (name, description, input schema) for one action in your business. Then list three validation or permission checks your code must perform before executing it.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.