Voice AI & Conversational AgentsPlatforms and stacks · Lesson 4 of 17
Building an agent with ElevenAgents
Video lecture
Building an agent with ElevenAgents
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Building with ElevenAgents
In this lesson, you'll go from blank page to a working voice agent on ElevenAgents, the ElevenLabs platform for voice and chat agents that used to be called Conversational AI. By the end, you'll know every major configuration area: the agent's prompt and language model, the voice, turn-taking, tools, knowledge base, workflows, testing and analysis. And you'll see how to create a tool and an agent through the API, so your setup can be versioned like code.
0:33 Why a platform
Why use an integrated platform at all, instead of wiring components yourself? Speed and completeness. The hard parts of voice, like turn-taking, interruptions, telephony, testing and analytics, are already solved and maintained. You focus on the conversation design, the tools and the business rules. For most businesses, that's the right trade-off until they have unusual needs or large engineering teams.
0:59 What's in the box
ElevenAgents bundles the whole cascade: speech recognition, a choice of language models from several providers or your own custom model endpoint, ElevenLabs voices, turn-taking, tools, retrieval over a knowledge base, workflows, testing and analytics. It deploys to a web widget, mobile apps, phone lines through Twilio or SIP trunking, WhatsApp and other channels. The core settings: the agent's system prompt, first message, language and model; the voice and its model; turn-taking behavior; tools; the knowledge base; analysis; and security and privacy controls.
1:35 Analogy: onboarding a receptionist
An analogy: setting up an agent on a platform is like onboarding a new receptionist. The system prompt is their job description. The first message is how they answer the phone. The voice is how they sound. Tools are the systems they're allowed to use, like the calendar and the transfer button. The knowledge base is the binder of policies on their desk. Tests are the role-plays you run before their first day. And evaluation criteria are how you review their calls each week.
2:12 Dynamic variables
Dynamic variables make agents personal without rewriting prompts. You write them in double curly braces, like customer name or clinic branch, and they're filled when a conversation starts, for example from your CRM when placing an outbound call, or updated by tool responses. System variables, like the current time and caller I D, are available too. So one prompt can greet Fatima by name at the Dubai Marina branch and James at the London branch.
2:45 Tools
Tools are how agents do real work. Server tools, also called webhooks, call your HTTPS endpoint with parameters the model extracts from the conversation, like check availability with a date and service. You describe the tool and its parameters, and your server returns JSON. You control whether the agent speaks before the tool runs, the timeout, and whether callers can interrupt meanwhile. Client tools trigger things in the user's app, like showing a calendar. And system tools provide built-in behaviors: ending the call, detecting language, transferring to a human or another agent, and detecting voicemail.
3:26 Structure for reliability
For complex calls, split responsibilities. A triage agent can hand off to a booking agent or a billing agent, or you can build a visual workflow with nodes and conditions. Smaller, focused prompts are easier to test and more reliable than one giant prompt. And put reference material, like your price list, opening hours and policies, into the knowledge base with retrieval turned on, instead of stuffing it all into the prompt. That keeps the prompt lean, which also helps latency.
4:01 Test and analyze
Testing is built in. Response tests check whether a reply meets a success condition. Tool-call tests check the agent called the right tool with the right parameters. And simulation tests let a simulated caller, with a scenario and persona, talk to your agent for several turns, judged against success conditions, with tools mocked if needed. After launch, evaluation criteria and data collection run on every conversation, and a post-call webhook can send transcripts and results to your systems. Verify that webhook's signature with your secret.
4:38 Worked example: Aria for Nova Dental
Here's Aria, a dental receptionist for Nova Dental in Dubai and London. Her first message says she's Nova Dental's virtual assistant and what she can do. She uses a calm designed voice on the Flash model. Her tools check availability, book and cancel, plus system tools to transfer to reception, end the call and detect language. Her knowledge base holds hours, services, prices and policies. Her prompt forbids clinical advice and requires confirming date, time and branch before booking. And every call is scored: booked or changed, no clinical advice, correct disclosure.
5:18 Example 2: Abu Dhabi salon
A simple example. A salon in Abu Dhabi builds its first agent in the dashboard, not the API. It answers questions about services and prices from the knowledge base, collects a name and preferred time, and uses one webhook tool that emails the request to the front desk, instead of booking directly. A system tool transfers to the salon's number during opening hours. It's not fully automated, but it captures every after-hours enquiry, and the team can add direct booking later.
5:53 Common mistakes
Common mistakes on agent platforms. One giant prompt that tries to handle every scenario, instead of focused agents or workflow nodes. Stuffing the full price list or policies into the prompt instead of the knowledge base. Launching without any tests, so you can't tell whether a change helped or hurt. And forgetting to verify the post-call webhook signature, which lets anyone send fake call results into your systems.
6:23 Watch me do it: create tool + agent
Watch me do it. I run the API script from the lesson text step by step. First, the environment: my ElevenLabs key and the booking API URL are in environment variables, never in the file. I run the tool creation call for check availability. The response comes back with a tool I D, and I paste it into the config. Next, I open the system prompt file, aria system prompt dot M D, and check the first line of the process section: call check availability before offering times. Then I run the agent creation call. It returns an agent I D. I open the agent in the dashboard to confirm everything landed: the first message says Aria is Nova Dental's virtual assistant, the voice is the one we designed, the Flash model is selected, and the tool appears under tools. I add two system tools in the dashboard: transfer to number, pointing to reception, with the condition caller asks for a person or is upset, and end call. I upload the price list and opening hours as knowledge base documents. Then, before any test call, I create three simulation tests: a simple cleaning booking, a caller who changes the date, and a caller asking whether a tooth pain is serious. I run them. Two pass. The third fails: the agent gave a guess about the pain. I tighten the guardrail line and re-run until it passes.
8:06 Recap and next step
Recap. ElevenAgents covers the full stack, from voice to analytics. Use dynamic variables for personalization, webhook tools for real actions, system tools for transfers and call control, the knowledge base for reference content, and tests plus evaluation criteria so you can tell whether changes help. Your next step: follow the API example in the lesson text to create one webhook tool and one agent, or build the same in the dashboard, then write three simulation tests before you make a single test call.
What ElevenAgents gives you
ElevenAgents is ElevenLabs' platform for voice and chat agents (it was previously called "Conversational AI", and some API paths still use convai). It bundles the whole cascade: speech recognition, a choice of LLMs (from several providers, or your own via a custom LLM endpoint), ElevenLabs TTS voices, turn-taking, tools, a knowledge base with retrieval, workflows, testing, analytics and deployment to web, mobile, phone (Twilio, SIP trunking), WhatsApp and other channels. Features evolve quickly; check the docs for current names and limits.
The core configuration
| Area | What you set |
|---|---|
| Agent | System prompt, first message, default language, LLM choice and parameters (temperature, max tokens), dynamic variables |
| Voice (TTS) | Voice, conversational TTS model, stability/speed/similarity, pronunciation dictionaries |
| Turn-taking | Turn timeout, eagerness (patient/normal/eager), silence end-call timeout, interruption behavior |
| Tools | Server (webhook) tools, client tools, MCP servers, and system tools such as end call, language detection, transfer to another agent, transfer to a phone number, skip turn, play keypad tones (DTMF) and voicemail detection |
| Knowledge base | Upload documents, URLs or text; enable retrieval (RAG) so the agent pulls relevant chunks instead of stuffing everything into the prompt |
| Analysis | Evaluation criteria (did the call achieve its goal?) and data collection fields (extract structured values from each conversation) |
| Security and privacy | Authentication for widgets, allowlists, overrides, data retention, guardrails |
Dynamic variables
Variables in double curly braces, like {{customer_name}} or {{clinic_branch}}, are filled at conversation start (for example from your CRM when placing an outbound call) or updated by tool responses. System variables such as the current time and caller ID are also available. Use them to personalize without rewriting prompts.
Tools: the three kinds you will use most
- Server tools (webhooks): the platform calls your HTTPS endpoint with parameters the LLM extracts (for example
check_availability(date, service)). You describe the tool and its JSON parameters; your server does the work and returns JSON. You can control whether the agent speaks before the tool runs ("pre-tool speech"), timeouts, and whether callers can interrupt during the tool. - Client tools: trigger actions in the user's app or web page (show a calendar, open a product page).
- System tools: built-in behaviors such as ending the call, detecting language, transferring to a human number or another agent, and detecting voicemail.
Workflows and multi-agent handoffs
For complex calls, split responsibilities: a triage agent that transfers to a booking agent or a billing agent, or a visual workflow with nodes and conditions. Smaller, focused prompts are easier to test and usually more reliable than one giant prompt.
Testing and analysis
ElevenAgents supports automated tests: response tests (does the agent's reply meet a success condition?), tool-call tests (did it call the right tool with the right parameters?) and simulation tests where a simulated caller with a scenario and persona talks to your agent for several turns and is judged against success conditions, with the option to mock tools. After launch, evaluation criteria and data collection run on every conversation, and a post-call webhook can send transcripts and results to your systems (verify the signature with your webhook secret).
Worked example: "Aria" for Nova Dental (Dubai and London)
- Agent: first message "Hi, this is Aria, Nova Dental's virtual assistant. I can book, move or cancel appointments. How can I help?" Language English with Arabic available.
- Voice: designed calm voice; Flash model; slightly slow speed.
- Tools:
check_availability,book_appointment,cancel_appointment(webhooks); system tools: transfer to reception number, end call, language detection. - Knowledge base: opening hours, services, price list, preparation instructions, cancellation policy.
- Guardrails in prompt: never give clinical advice; emergencies go to emergency services wording; always confirm date, time and branch before booking.
- Evaluation criteria: "appointment_booked_or_changed", "no_clinical_advice", "correct_disclosure". Data collection: service requested, branch, new-patient flag.
Hands-on: create a webhook tool and an agent via the API
import os, requests
API = "https://api.elevenlabs.io/v1/convai"
HEADERS = {"xi-api-key": os.environ["ELEVENLABS_API_KEY"], "Content-Type": "application/json"}
tool = {
"tool_config": {
"type": "webhook",
"name": "check_availability",
"description": "Check free appointment slots for a service at a branch on a given date. Use before offering times.",
"api_schema": {
"url": os.environ["BOOKING_API_URL"] + "/availability",
"method": "POST",
"request_body_schema": {
"type": "object",
"properties": {
"date": {"type": "string", "description": "Requested date in YYYY-MM-DD"},
"service": {"type": "string", "description": "Service name, e.g. cleaning, check-up"},
"branch": {"type": "string", "description": "Branch: dubai-marina or london-soho"}
},
"required": ["date", "service", "branch"]
}
}
}
}
r = requests.post(f"{API}/tools", json=tool, headers=HEADERS, timeout=30)
r.raise_for_status()
tool_id = r.json()["id"]
agent = {
"name": "Aria - Nova Dental",
"conversation_config": {
"agent": {
"first_message": "Hi, this is Aria, Nova Dental's virtual assistant. How can I help today?",
"language": "en",
"prompt": {
"prompt": open("aria_system_prompt.md", encoding="utf-8").read(),
"llm": os.environ.get("AGENT_LLM", "gemini-2.5-flash"),
"temperature": 0.3,
"tool_ids": [tool_id]
}
},
"tts": {"voice_id": os.environ["BRAND_VOICE_ID"], "model_id": "eleven_flash_v2_5"},
"turn": {"turn_eagerness": "normal"}
},
"platform_settings": {
"evaluation": {"criteria": [
{"id": "booked", "name": "Appointment booked or changed",
"conversation_goal_prompt": "Did the agent successfully book, move or cancel an appointment the caller asked for?"}
]},
"data_collection": {
"service": {"type": "string", "description": "The dental service the caller wanted"}
}
}
}
r = requests.post(f"{API}/agents/create", json=agent, headers=HEADERS, timeout=30)
r.raise_for_status()
print("agent_id:", r.json()["agent_id"])The field names above follow the public API at the time of writing; confirm against the current API reference, and choose the LLM from the platform's supported list. Many teams configure agents in the dashboard first, then export and version the configuration.
Pitfalls
- One giant prompt that tries to handle every scenario; split into focused agents or workflow nodes.
- Stuffing the whole price list into the prompt instead of the knowledge base.
- Launching without tests or evaluation criteria, so you cannot tell whether changes help or hurt.
Key takeaways
- ElevenAgents (formerly Conversational AI) bundles ASR, LLM choice, TTS, turn-taking, tools, knowledge base, workflows, tests and analytics across web, phone and messaging channels.
- Dynamic variables personalize conversations from CRM data or tool results without rewriting prompts.
- Use webhook tools for actions, system tools for transfers, end call, language detection and voicemail, and the knowledge base for reference content.
- Write response, tool-call and simulation tests and define evaluation criteria and data collection before launch.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Create one webhook tool and one agent (via API or dashboard) for a booking use case. Before any test call, write three simulation tests with clear success conditions.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.