Latest AI Techniques: RAG, Tool Use, Agents & MCPTool and function calling · Lesson 6 of 20
Function calling fundamentals
Video lecture
Function calling fundamentals
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Tool use (function calling)
A language model on its own can only produce text. It can't check a stock level, look up an order or book a meeting. Tool use, often called function calling, changes that. The model asks your application to run a function, your code runs it, and the result goes back to the model. In this lesson you'll learn the request cycle, how to define a tool the model will use correctly, and why your code, not the prompt, must be the security boundary.
0:36 Analogy: the manager and the assistant
Here's the analogy. Think of the model as a brilliant manager who can't touch the keyboard. It can read, reason and decide what needs doing, but it has to ask an assistant, your code, to actually look things up or press buttons. A good assistant checks the request makes sense and that the person asking is allowed, before doing anything. That's the whole shape of tool use: the model asks, your code checks and acts.
1:09 The cycle
Here's the cycle. You send the conversation plus a list of tool definitions: a name, a description and a JSON schema for the inputs. The model replies either with a normal answer or with a tool call, such as get order status with order ID A20931. Your code validates the arguments, checks permissions and runs the function. You send back the result linked to that call. The model then answers, or calls another tool. The model never executes anything itself, apart from special hosted tools that run on the provider's side.
1:49 Descriptions are prompts
Descriptions are prompts. The model decides whether and how to call a tool mostly from its name and description. A great description says what the tool does and returns, when to use it and when not to, what each input means with formats and examples, and its limits, like returns at most twenty results. Compare gets data with, look up the status and delivery date of one order belonging to the signed-in customer; if no order ID was given, ask for it rather than guessing. The second one prevents a whole class of mistakes.
2:30 Simple example: a forecast tool
A simple example. A customer types: what's the weather in Dubai tomorrow? You've given the model one tool, get forecast, with a city and a date. The model replies not with text, but with a structured request: get forecast, city Dubai, date tomorrow's date. Your code calls a weather service, gets back thirty-eight degrees and sunny, and sends that result back. The model then writes: tomorrow in Dubai looks sunny, around thirty-eight degrees. Four messages, one tool, and live information the model could never have memorised.
3:07 Your code is the boundary
Now security. A tool call is a request, and requests can be wrong or manipulated, for example by instructions hidden in an email the model read earlier. So your code validates arguments against the schema and business rules, authorises using the real user's identity, never the model's claims, asks for confirmation on consequential actions like refunds or sending messages, applies rate limits and timeouts, and logs every call. If a customer asks about someone else's order, the permission check fails in code, and the model simply explains politely.
3:45 Return results well
How you return results matters too. Keep them concise and structured. A five-thousand-line API response wastes context and distracts the model, so extract what matters. Return errors that teach, like order not found, ask the customer to check the ID in their confirmation email, and mark them as errors so the model knows the call failed. And flag untrusted content, such as text from web pages or emails, so the model treats it as data rather than instructions.
4:19 Failure modes
Watch for three classic failure modes. Hallucinated arguments, where the model invents an order ID the user never gave, which you counter with schema patterns, a description that says ask if missing, and validation. Tool overload, where dozens of similar tools confuse selection, which you counter by consolidating or loading tools dynamically. And silent failures, where a tool returns empty results and the model answers as if it had data, which you counter with explicit no-results messages.
4:52 Common mistakes
Common mistakes, beyond the failure modes you just heard. Putting authorisation rules only in the system prompt. Returning raw database rows full of fields the model doesn't need. Swallowing errors so the model thinks a failed call succeeded. Giving tools vague names like process or handle. And forgetting to limit the loop, so a confused model can call tools forever. Each of these is cheap to prevent and expensive to discover in production.
5:24 How you'll know it's working
How will you know your tools are working? Track tool-call success rate, meaning calls that pass validation and execute. Track errors by type: invalid arguments, not authorised, not found. Watch how often the model asks the user for missing information instead of guessing. And review a sample of conversations weekly for calls that should or shouldn't have happened. Falling invalid-argument errors usually mean your descriptions are getting better.
5:54 Business example (illustrative)
A deeper business example, illustrative. A Lahore e-commerce brand's support chat used to answer where is my order with a generic tracking link, and customers often followed up with a phone call. After adding an order-status tool that checks ownership and returns status and expected delivery, around two thirds of those chats ended without a call. The team also noticed attempted lookups of other people's orders in the logs, all blocked by the ownership check in code.
6:27 Hands-on in the lesson
In the lesson you'll find a complete, working tool loop in Python using the Claude API. It defines an order-status tool with a strict schema, validates the ID, checks that the order belongs to the signed-in customer, returns errors with the error flag set, sends all results back in a single message, and stops after a fixed number of turns. Run it twice: once as the right customer, once as the wrong one, and watch the authorisation path. Other platforms use different field names for the same cycle, so once you understand this one, you can read them all.
7:10 Watch me do it: the tool loop
Watch me do it. I open the tool loop. First, the tool definition: a name, a description that says ask for the ID rather than guessing, and a schema with a pattern for IDs. Next, get order status validates the ID, then checks that the order's owner matches the customer ID from the session, not from the model. Then run starts the messages list and loops up to six turns. Each turn, I call the API with the tools and append the assistant's content. If the stop reason isn't tool use, I return the text. Otherwise, for each tool use block, I try the function, and on success add a tool result with the JSON. On an error, I add a tool result with the error message and is error set to true. All results go back in one user message. I run it twice. As customer seventeen, I get the delivery date. As customer ninety-nine, the model politely says it can't find that order on this account.
8:23 Recap
To recap: the model requests tool calls, and your code validates, authorises, executes and returns results. Names, descriptions and schemas are prompts, so write them carefully. Enforce permissions in code, return concise results and helpful errors, and treat outside content as untrusted. Your next step is to write one tool definition for an action in your business, then list three checks your code must run before executing it. Next, we'll look at designing tools that models use reliably.
8:57 Try this now (15 minutes)
Try this now. Choose one action in your business that an assistant could usefully do, like checking stock, looking up a booking or creating a draft invoice. Write its tool definition: a clear name, a description saying what it does, when to use it and when not to, and an input schema with formats. Then list three checks your code must run before executing it. If you have ten more minutes, plug it into the hands-on loop with fake data.
What tool use is
Tool use (also called function calling) lets a model request that your application run a function: search a database, call an API, do a calculation, send a message. The model does not execute anything itself. It outputs a structured request (the tool name and arguments), your code executes it, and you return the result to the model, which then continues.
This turns a model from a text generator into the reasoning layer of a system that can fetch live data and take actions.
The request cycle
1. You send: messages + tool definitions (name, description, input schema)
2. Model replies: either a normal answer, or a tool call such as
get_order_status({"order_id": "A1234"})
3. Your code: validates arguments, checks permissions, executes the function
4. You send back: the tool result, linked to that call
5. Model replies: a final answer using the result, or another tool callMany APIs allow several tool calls in one turn (for example fetching three records in parallel), and let you control whether the model must use a tool, may use one, or must use a specific one.
Defining a tool
A tool definition is a name, a natural-language description and a JSON Schema for its inputs:
{
"name": "get_order_status",
"description": "Look up the current status and estimated delivery date of a customer's order. Use when the customer asks where their order is or when it will arrive. Requires the order ID, which starts with 'A' followed by 4-8 digits.",
"input_schema": {
"type": "object",
"properties": {
"order_id": {"type": "string", "pattern": "^A[0-9]{4,8}$",
"description": "Order ID from the confirmation email, e.g. A1234"}
},
"required": ["order_id"]
}
}The exact field names differ slightly between providers, but the ingredients are the same everywhere.
Descriptions are prompts
The model decides whether and how to call a tool mainly from its name and description. Good descriptions state:
- what the tool does and returns
- when to use it (and when not to)
- what inputs mean, with formats and examples
- limitations ("returns at most 20 results", "data refreshed hourly")
Vague descriptions ("Gets data") cause missed calls, wrong calls and invented arguments.
Executing safely
Your code is the security boundary. The model's tool call is a request, and it may be wrong or manipulated (for example via prompt injection in earlier content). So:
- Validate arguments against the schema and business rules.
- Authorise using the real user's permissions, not the model's say-so. If a user can only see their own orders, enforce that in code.
- Confirm consequential actions with the user (refunds, sending messages, deleting data).
- Rate-limit and time-out external calls.
- Log every call and result.
Returning results well
- Return concise, structured results. A 5,000-line API response wastes context and distracts the model; extract what matters.
- Return useful errors: "Order A1234 not found. Ask the customer to check the ID in their confirmation email" helps the model recover. A raw stack trace does not.
- Mark untrusted content (for example text from web pages or emails) so the model treats it as data.
Worked example: an order-tracking assistant
User: Where's my order? It's A20931.
Model -> get_order_status({"order_id": "A20931"})
App -> checks that A20931 belongs to the logged-in customer; calls the
order system; returns {"status": "in_transit", "carrier": "...",
"eta": "2026-09-29"}
Model -> "Your order is on its way and is expected on 29 September..."If the customer asks about someone else's order, the app's permission check fails and returns "Not authorised for this order", and the model explains politely. The model never had the power to bypass that check.
Failure modes
- Hallucinated arguments: the model invents an order ID when the user didn't provide one. Mitigate with schema patterns, descriptions ("ask the user if missing") and validation.
- Tool overload: dozens of similar tools confuse selection. Consolidate, or load tools dynamically based on the task.
- Silent failures: a tool returns empty results and the model answers as if it had data. Return explicit "no results" messages.
Hands-on: a complete tool-use loop with validation and authorisation
Below is a manual tool loop with the Claude API (Python SDK). It shows every step you read about: sending a tool definition, receiving a tool_use block, validating and authorising in your own code, returning a tool_result (with is_error when something fails) and letting the model finish. The SDK also offers a beta "tool runner" helper that drives this loop for you; writing it once by hand is the best way to understand what such helpers do.
import json, os, re
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5") # check current models in the docs
TOOLS = [{
"name": "get_order_status",
"description": ("Look up status and estimated delivery of ONE order belonging to the signed-in "
"customer. Use when they ask where an order is. If no order ID was given, ask "
"for it instead of guessing. IDs look like A1234 (A + 4-8 digits)."),
"input_schema": {"type": "object",
"properties": {"order_id": {"type": "string", "pattern": "^A[0-9]{4,8}$"}},
"required": ["order_id"], "additionalProperties": False},
}]
FAKE_DB = {"A20931": {"owner": "cust_17", "status": "in_transit", "eta": "2026-09-29"}}
def get_order_status(order_id: str, customer_id: str) -> dict:
if not re.fullmatch(r"A[0-9]{4,8}", order_id):
raise ValueError("order_id must look like A1234; ask the customer to check their email")
order = FAKE_DB.get(order_id)
if order is None or order["owner"] != customer_id: # authorise in code, not in the prompt
raise PermissionError("No order with that ID for this customer")
return {"order_id": order_id, "status": order["status"], "eta": order["eta"]}
def run(user_text: str, customer_id: str, max_turns: int = 6) -> str:
messages = [{"role": "user", "content": user_text}]
for _ in range(max_turns):
resp = client.messages.create(model=MODEL, max_tokens=4000, tools=TOOLS,
system="You are a concise order-support assistant.",
messages=messages)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason != "tool_use":
return "".join(b.text for b in resp.content if b.type == "text")
results = []
for block in resp.content:
if block.type != "tool_use":
continue
try:
out = get_order_status(block.input["order_id"], customer_id)
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": json.dumps(out)})
except (ValueError, PermissionError, KeyError) as err:
results.append({"type": "tool_result", "tool_use_id": block.id,
"content": str(err), "is_error": True})
messages.append({"role": "user", "content": results}) # all results in ONE message
return "Sorry, I could not complete that. A colleague will follow up."
print(run("Where's my order A20931?", customer_id="cust_17"))
print(run("Where's order A20931?", customer_id="cust_99")) # authorisation failure pathNotice three production habits: the permission check uses the real signed-in customer, not anything the model says; errors are returned as readable tool_result content with is_error so the model can recover gracefully; and the loop has a turn limit.
The same idea on other platforms
OpenAI's Responses API, Google's Gemini API and open-source runtimes all implement the same cycle with slightly different field names (for example function_call and function_call_output items rather than tool_use and tool_result blocks). Frameworks such as the OpenAI Agents SDK wrap it in helpers. Learn the cycle once and every SDK becomes readable. The flagship course Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API covers provider differences, streaming tool calls and production error handling in depth.
Going further
Some providers offer server-side tools (web search, web fetch, code execution) that run in the provider's environment, alongside client-side tools you execute. The design principles are the same: clear descriptions, least privilege, and treating every result as input to be checked.
Key takeaways
- The model requests tool calls; your code validates, authorises, executes and returns results.
- Tool names, descriptions and schemas are prompts: state what, when, input formats and limits.
- Your code is the security boundary: validate arguments, use the real user's permissions, confirm consequential actions.
- Return concise, structured results and helpful errors; flag untrusted content.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Write a tool definition (name, description, input schema) for one action in your business. Then list three validation or permission checks your code must perform before executing it.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.