Voice AI & Conversational AgentsDesigning voice conversations · Lesson 9 of 17
Tools and knowledge in voice agents
Video lecture
Tools and knowledge in voice agents
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Tools and knowledge for voice
In a chat app, a four-second tool call is just a spinner. In a voice call, it's four seconds of dead air, and callers start saying hello, are you there? Tools and knowledge in voice agents have to be fast, forgiving and safe, and the agent has to talk naturally around them. In this lesson, you'll learn to design voice-friendly tools, talk through waits, build a knowledge base that retrieves well for speech, use MCP safely, and write a fast webhook.
0:35 Why tools + knowledge
Why focus on tools and knowledge? Because that's where voice agents create real value, by checking availability, booking, looking up orders and answering policy questions correctly. It's also where most production problems come from: slow APIs, confusing tool descriptions, stale knowledge and unsafe write actions. Getting this layer right is the difference between a pleasant talker and a useful assistant.
1:01 Six tool principles
Six principles for tools. One: keep them narrow and well described. Check availability with a date, service and branch beats a generic call A P I tool. Two: accept forgiving inputs. Callers say next Thursday, so give the model today's date or resolve dates on your server, and validate everything. Three: respond fast, ideally well under a second, with caching and pre-fetching at call start. Four: return speakable results, the best two or three options, not two hundred. Five: fail honestly with a message that helps the agent recover. Six: keep write actions safe with confirmation, identity checks and logs.
1:44 Analogy: hotel reception
An analogy: tools for a voice agent are like the phone system at a busy hotel reception. The receptionist can check room availability in seconds, make a booking only after the guest confirms, and call the manager when needed. If the booking system is slow, they say one moment while I check, rather than staring in silence. And they never promise a room the system says is taken. Fast lookups, confirmed writes, polite waits and honest answers.
2:17 Talking around tools
Now, talking around tools. Pre-tool speech, like let me check Thursday for you, covers the wait, and many platforms can insert it automatically when tools are slow. Subtle tool-call sounds, like typing, can fill longer waits, but use them sparingly. For long operations, like emailing a quote, run the tool asynchronously and tell the caller it's on its way. And decide on interruptions: usually you disable them while a booking is being written, then confirm the result.
2:50 Knowledge for speech
Voice agents answer policy and product questions from a knowledge base using retrieval. For speech, chunk content into short, self-contained pieces, like cancellation policy: cancel at least twenty-four hours before. Write it FAQ-style, a question with a one or two sentence spoken answer, then details. Keep it current, with an owner and review date, and re-index when prices or hours change. Don't retrieve too many chunks, because long context slows the model and invites rambling. And when nothing relevant is found, the agent should say it doesn't know and offer a transfer.
3:30 MCP in voice
The Model Context Protocol lets agents use tools exposed by MCP servers, like a calendar, CRM or ticketing system. ElevenAgents and OpenAI's Realtime API both support remote MCP servers. The benefit is reuse across agents. The cautions: MCP tools can be powerful, so use approval policies, for example requiring approval for all tools or specific ones, restrict which tools are exposed, and prefer read-only tools in voice. And remember each remote call adds network time.
4:03 Worked example: Lahore restaurants
A restaurant chain in Lahore built a reservation and order-status agent. Tools: find tables, create a reservation, which needs confirmation, and order status by order number or phone. Knowledge: menu highlights, halal certification, delivery zones and hours, in Urdu and English. At call start, the caller's phone number pre-fetches any active order, so status is instant. Order status returns one line, like out for delivery, rider ten minutes away. Reservation writes block interruptions until confirmed. And when retrieval confidence is low, the agent offers to connect the branch.
4:41 Hands-on: voice-friendly webhook
The lesson text includes a small FastAPI webhook for checking availability. It checks a shared secret header, validates the branch and date, rejects past dates, returns at most two slots, and returns a say message the agent can use when something's wrong, instead of an error the agent can't explain. Configure the secret as a header in the tool definition, stored as a platform secret, host the endpoint in the same region as your agent, and time it under load.
5:16 Example 2: Dubai car rental
A simple example. A Dubai car rental company's agent answers where's my booking. At call start, the caller's number pre-fetches their active reservation. The tool returns one line: pickup at the airport branch tomorrow at ten a m, compact car. The agent reads that back in one sentence, and offers to change the time. No search, no waiting, no list. Designing the tool for the voice answer, not for a web page, is what makes it feel instant.
5:50 Common mistakes
Common mistakes with tools and knowledge. Generic tools with vague descriptions, which lead the model to call the wrong tool or none at all. Returning long lists that the agent then reads aloud. Letting the agent answer anyway when a tool fails, which invents information. And knowledge bases with no owner, so prices and opening hours go stale and the agent confidently gives old answers.
6:18 Watch me do it: harden tool + knowledge
Watch me do it. I open the FastAPI availability webhook from the lesson text and run it locally. First, I call it with curl using a past date: it returns ok false with a say message asking for a future date. Then an unknown branch: it returns the list of valid branches in a say message. Then a real date: it returns exactly two slots, not the forty the database holds. Next, security: I call it without the shared-secret header and get a four oh one. Good. I deploy it in the same region as the agent platform and time twenty calls: the median is well under half a second. In the agent's tool definition I add the header with the secret stored as a platform secret, set the timeout, and switch on pre-tool speech. Now the knowledge base: I rewrite the price list as FAQ-style chunks, one question and a one or two sentence spoken answer each, in English and Arabic. I delete an old PDF of last year's prices, which was causing wrong answers. Then I enable retrieval and ask a test question: how much is whitening? The agent answers from the new chunk in one sentence. I ask about a service we don't offer, and it says it doesn't have that information and offers reception. That abstention is exactly what we want.
7:56 Recap and next step
Recap. Make tools narrow, fast, forgiving and safe. Talk around waits, and never let the agent invent results when a tool fails. Build speech-friendly knowledge with short chunks, owners and abstention. Use MCP with approvals and least privilege. Your next step: time every tool your agent uses on ten realistic requests. Anything slower than a second gets caching, pre-fetching or pre-tool speech, and any tool returning more than three options gets trimmed.
Voice raises the bar for tools
In chat, a 4-second tool call is a spinner. In voice, it is dead air. Tools and knowledge retrieval must be fast, forgiving and safe, and the agent must talk naturally around them.
Designing tools for voice
1. Narrow, well-described tools. check_availability(date, service, branch) beats a generic call_api(endpoint, payload). The description tells the LLM exactly when to use it; parameter descriptions include formats ("YYYY-MM-DD") and allowed values.
2. Forgiving inputs. Callers say "next Thursday" or "tomorrow afternoon". Either give the LLM today's date and time zone so it resolves dates, or accept natural language and resolve it server-side. Validate everything.
3. Fast responses. Aim for tools that respond well under a second. Cache reference data (branches, services, prices). Pre-fetch likely data at call start (for example the caller's upcoming appointments by caller ID).
4. Speakable results. Return concise JSON the LLM can turn into one sentence. Do not return 200 slots; return the best two or three. Filter responses to what the agent needs (some platforms support response filters).
5. Honest failure. On timeout or error, return a message that lets the agent recover: "Availability service unavailable. Offer to take a callback number." Never let the agent invent results.
6. Safe actions. Separate read tools from write tools. Require explicit caller confirmation before write actions. Enforce identity and business rules in the API. Log every action.
Talking around tools
- Pre-tool speech: "Let me check Thursday for you." Many platforms can insert this automatically when tools are slow.
- Tool-call sounds: subtle typing or hold sounds can fill longer waits (use sparingly).
- Async tools: for long operations (sending a PDF quote by email), run in the background and tell the caller it's on its way.
- Interruptions during tools: decide whether the caller can interrupt while a booking is being written; usually disable interruption during the write, then confirm.
Knowledge: retrieval for voice
Voice agents answer policy and product questions from a knowledge base with retrieval-augmented generation (RAG):
- Chunk for speech: short, self-contained chunks ("Cancellation policy: cancel at least 24 hours before...") retrieve better and produce shorter answers.
- Write FAQ-style content: question plus a one- or two-sentence spoken answer, plus details.
- Keep it current: pricing and hours change; assign an owner and review date; re-index on change.
- Limit context: retrieving too many chunks slows the LLM and invites rambling.
- Abstain: instruct the agent to say it doesn't know and offer a transfer when retrieval finds nothing relevant.
MCP in voice agents
The Model Context Protocol (MCP) lets agents connect to tools exposed by MCP servers (for example a calendar, CRM or ticketing system). ElevenAgents and OpenAI's Realtime API support remote MCP servers. Benefits: reuse integrations across agents. Cautions: MCP tools can be powerful; use approval policies (for example require approval for all or specific tools), restrict which tools are exposed, and prefer read-only tools for voice where possible. Latency matters: remote MCP calls add network hops.
Worked example: a Lahore restaurant chain's reservation and order-status agent
Tools: find_tables(branch, datetime, party_size), create_reservation(...) (write, confirmation required), order_status(order_id or phone) (read). Knowledge: menu highlights, halal certification statement, delivery zones, hours per branch, in Urdu and English.
Design decisions: caller ID prefetch finds active orders instantly; order_status returns a one-line status ("Out for delivery, rider 10 minutes away"); reservation writes disable interruptions until confirmed; knowledge chunks are FAQ-style in both languages; if retrieval confidence is low, the agent offers to connect the branch.
Hands-on: a fast, voice-friendly webhook (FastAPI)
import os, datetime as dt
from fastapi import FastAPI, Header, HTTPException
from pydantic import BaseModel
app = FastAPI()
TOOL_SECRET = os.environ["TOOL_SHARED_SECRET"] # also configured as a header in the agent tool
class AvailabilityIn(BaseModel):
date: str # YYYY-MM-DD
service: str
branch: str
BRANCHES = {"dubai-marina", "downtown-dubai", "london-soho"}
@app.post("/availability")
def availability(body: AvailabilityIn, x_tool_secret: str = Header(default="")):
if x_tool_secret != TOOL_SECRET:
raise HTTPException(status_code=401, detail="unauthorized")
if body.branch not in BRANCHES:
return {"ok": False, "say": "That branch isn't one of ours. Offer Dubai Marina, Downtown Dubai or London Soho."}
try:
day = dt.date.fromisoformat(body.date)
except ValueError:
return {"ok": False, "say": "Ask the caller to confirm the date."}
if day < dt.date.today():
return {"ok": False, "say": "That date has passed. Ask for a future date."}
slots = find_slots(day, body.service, body.branch) # your cached lookup; keep it fast
if not slots:
return {"ok": True, "slots": [], "say": "No slots that day. Offer the next available day."}
return {"ok": True, "slots": slots[:2]} # at most two options for voiceConfigure the shared secret as a request header in the tool definition (stored as a platform secret), keep the endpoint in the same region as the agent, and time it.
Pitfalls
- Generic tools with vague descriptions, leading to wrong or missing calls.
- Returning long lists that the agent reads aloud.
- Letting the agent "answer anyway" when a tool fails.
Key takeaways
- Voice tools must be narrow, well described, fast, forgiving of natural inputs, strictly validated and safe for writes.
- Return speakable results (two or three options) and honest failure messages the agent can use to recover.
- Use pre-tool speech, async tools for long jobs and no interruptions during writes.
- Build knowledge bases of short FAQ-style chunks with owners, and use MCP with approval policies and least privilege.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Time every tool your agent uses on ten realistic requests. Add caching, pre-fetching or pre-tool speech to anything slower than a second, and trim results to at most three options.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.