Voice AI & Conversational AgentsDesigning voice conversations · Lesson 9 of 17

Tools and knowledge in voice agents

Article · 16 min · 8 min lecture

Video lecture

Tools and knowledge in voice agents

13 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 13

Tools and knowledge for voice

  • Dead air kills calls
  • Fast, forgiving, safe tools
  • Speech-friendly knowledge
  • MCP with care

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Voice raises the bar for tools

In chat, a 4-second tool call is a spinner. In voice, it is dead air. Tools and knowledge retrieval must be fast, forgiving and safe, and the agent must talk naturally around them.

Designing tools for voice

1. Narrow, well-described tools. check_availability(date, service, branch) beats a generic call_api(endpoint, payload). The description tells the LLM exactly when to use it; parameter descriptions include formats ("YYYY-MM-DD") and allowed values.

2. Forgiving inputs. Callers say "next Thursday" or "tomorrow afternoon". Either give the LLM today's date and time zone so it resolves dates, or accept natural language and resolve it server-side. Validate everything.

3. Fast responses. Aim for tools that respond well under a second. Cache reference data (branches, services, prices). Pre-fetch likely data at call start (for example the caller's upcoming appointments by caller ID).

4. Speakable results. Return concise JSON the LLM can turn into one sentence. Do not return 200 slots; return the best two or three. Filter responses to what the agent needs (some platforms support response filters).

5. Honest failure. On timeout or error, return a message that lets the agent recover: "Availability service unavailable. Offer to take a callback number." Never let the agent invent results.

6. Safe actions. Separate read tools from write tools. Require explicit caller confirmation before write actions. Enforce identity and business rules in the API. Log every action.

Talking around tools

  • Pre-tool speech: "Let me check Thursday for you." Many platforms can insert this automatically when tools are slow.
  • Tool-call sounds: subtle typing or hold sounds can fill longer waits (use sparingly).
  • Async tools: for long operations (sending a PDF quote by email), run in the background and tell the caller it's on its way.
  • Interruptions during tools: decide whether the caller can interrupt while a booking is being written; usually disable interruption during the write, then confirm.

Knowledge: retrieval for voice

Voice agents answer policy and product questions from a knowledge base with retrieval-augmented generation (RAG):

  • Chunk for speech: short, self-contained chunks ("Cancellation policy: cancel at least 24 hours before...") retrieve better and produce shorter answers.
  • Write FAQ-style content: question plus a one- or two-sentence spoken answer, plus details.
  • Keep it current: pricing and hours change; assign an owner and review date; re-index on change.
  • Limit context: retrieving too many chunks slows the LLM and invites rambling.
  • Abstain: instruct the agent to say it doesn't know and offer a transfer when retrieval finds nothing relevant.

MCP in voice agents

The Model Context Protocol (MCP) lets agents connect to tools exposed by MCP servers (for example a calendar, CRM or ticketing system). ElevenAgents and OpenAI's Realtime API support remote MCP servers. Benefits: reuse integrations across agents. Cautions: MCP tools can be powerful; use approval policies (for example require approval for all or specific tools), restrict which tools are exposed, and prefer read-only tools for voice where possible. Latency matters: remote MCP calls add network hops.

Worked example: a Lahore restaurant chain's reservation and order-status agent

Tools: find_tables(branch, datetime, party_size), create_reservation(...) (write, confirmation required), order_status(order_id or phone) (read). Knowledge: menu highlights, halal certification statement, delivery zones, hours per branch, in Urdu and English.

Design decisions: caller ID prefetch finds active orders instantly; order_status returns a one-line status ("Out for delivery, rider 10 minutes away"); reservation writes disable interruptions until confirmed; knowledge chunks are FAQ-style in both languages; if retrieval confidence is low, the agent offers to connect the branch.

Hands-on: a fast, voice-friendly webhook (FastAPI)

import os, datetime as dt
from fastapi import FastAPI, Header, HTTPException
from pydantic import BaseModel

app = FastAPI()
TOOL_SECRET = os.environ["TOOL_SHARED_SECRET"]  # also configured as a header in the agent tool

class AvailabilityIn(BaseModel):
    date: str      # YYYY-MM-DD
    service: str
    branch: str

BRANCHES = {"dubai-marina", "downtown-dubai", "london-soho"}

@app.post("/availability")
def availability(body: AvailabilityIn, x_tool_secret: str = Header(default="")):
    if x_tool_secret != TOOL_SECRET:
        raise HTTPException(status_code=401, detail="unauthorized")
    if body.branch not in BRANCHES:
        return {"ok": False, "say": "That branch isn't one of ours. Offer Dubai Marina, Downtown Dubai or London Soho."}
    try:
        day = dt.date.fromisoformat(body.date)
    except ValueError:
        return {"ok": False, "say": "Ask the caller to confirm the date."}
    if day < dt.date.today():
        return {"ok": False, "say": "That date has passed. Ask for a future date."}
    slots = find_slots(day, body.service, body.branch)  # your cached lookup; keep it fast
    if not slots:
        return {"ok": True, "slots": [], "say": "No slots that day. Offer the next available day."}
    return {"ok": True, "slots": slots[:2]}  # at most two options for voice

Configure the shared secret as a request header in the tool definition (stored as a platform secret), keep the endpoint in the same region as the agent, and time it.

Pitfalls

  • Generic tools with vague descriptions, leading to wrong or missing calls.
  • Returning long lists that the agent reads aloud.
  • Letting the agent "answer anyway" when a tool fails.

Key takeaways

  • Voice tools must be narrow, well described, fast, forgiving of natural inputs, strictly validated and safe for writes.
  • Return speakable results (two or three options) and honest failure messages the agent can use to recover.
  • Use pre-tool speech, async tools for long jobs and no interruptions during writes.
  • Build knowledge bases of short FAQ-style chunks with owners, and use MCP with approval policies and least privilege.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A tool returns 150 available slots. What is the voice-friendly fix?
  2. The booking API times out. What should the tool return?
  3. Which is a sensible MCP policy for a voice agent with access to a CRM?

Put it into practice

Time every tool your agent uses on ten realistic requests. Add caching, pre-fetching or pre-tool speech to anything slower than a second, and trim results to at most three options.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.