Voice AI & Conversational AgentsQuality, evaluation and operations · Lesson 14 of 17
Monitoring, cost control and scaling
Video lecture
Monitoring, cost control and scaling
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 After launch
Launch day isn't the finish line. Real callers are far more varied than your tests, models get updated, APIs slow down, and costs creep up quietly. In this lesson, you'll learn what to monitor, how to build feedback loops that make the agent better every week, how to roll out changes safely, how voice agent costs really work, and how to scale for peaks like Ramadan evenings or big sale days.
0:31 Why operations
Why does operations deserve its own lesson? Because the day you launch is the day your agent meets reality: new accents, unexpected questions, slow APIs, peak seasons and model updates. Without monitoring and a weekly improvement rhythm, quality quietly decays and costs quietly climb. With them, the agent gets better every week.
0:54 Seven signals
Monitor seven signals. Call volume and outcomes, for business health. Transfer and abandonment rates, for friction. Voice to voice latency, median and p90. Tool errors and timeouts. Evaluation criteria failures, including any disclosure or safety failure, which should alert immediately. Language mix and success per language, to catch hidden gaps. And cost per call and per successful outcome. Set alerts on changes, like success dropping compared with last week, not just on absolute numbers.
1:26 Analogy: running a call center
An analogy: running a voice agent in production is like running a small call center. You watch the dashboard for queues and dropped calls, you listen to a sample of calls each week, you coach on common mistakes, you staff up for busy seasons, and you know your cost per resolved call. The difference is that your coaching happens through prompts, tools and tests, and staffing up means capacity limits and load testing.
1:58 Feedback loops
Build feedback loops. Review a sample of transcripts every week, especially failures, transfers and low scores. Look at topics and intents: what are people calling about that the agent can't handle? Turn every production failure into a new test case, so it can never silently come back. And turn unanswered questions into knowledge base entries with owners. That weekly rhythm is what separates agents that improve from agents that decay.
2:28 Change management
Manage change carefully. Version everything: prompt, tools, knowledge, voice, model and turn settings. Roll out in stages, sending a slice of traffic to the new version, which some platforms support with branches and traffic splits, compare outcomes, then promote. Keep a one-click rollback. And track model deprecation notices from your providers, testing replacements early rather than scrambling at the deadline.
2:54 Understanding cost
Now cost. For a cascaded agent, cost per minute is roughly speech recognition, plus language model tokens, which grow with history and retrieved context, plus characters spoken by text to speech, plus telephony per-minute rates and number rental, plus any platform fee and tool costs. Speech to speech models bill audio tokens instead. Prices change often, so always use current pricing pages. And the number that matters isn't cost per minute. It's cost per successful outcome: per booking, per qualified lead, per resolved call, compared with the human alternative.
3:33 Cost levers and scaling
Six levers cut cost without hurting quality. Shorter agent turns, which also improve experience. Lean context, with trimmed history, fewer retrieved chunks and prompt caching where available. Right-sized models: a fast model for routine turns, escalating only when needed. Clean endings, with silence timeouts and maximum call durations. Voicemail detection, so you don't pay to talk to machines. And telephony choices, like regional carriers and SIP trunks. Scaling is the other side: know your concurrency limits, load test your tools, keep human backup for spikes, and deploy near your callers.
4:12 Worked example: KSA sale day
A Saudi e-commerce brand prepared its order-status line for a big sale. Before: they load-tested the order API at triple normal concurrency, raised concurrency limits, cached a sale FAQ with delivery times and returns, set a maximum call duration, and added a text message link for detailed tracking to keep calls short. During the sale, latency crept up as the order API slowed, so they turned on pre-tool speech and a cached status fallback. Afterwards, they turned the ten most common failures into tests and new FAQ entries. Cost per resolved call fell, because calls were shorter and more were resolved.
4:56 Example 2: UK outage line
A simple example of cost control. A UK utility's outage line agent averaged long calls because it read the full outage update every time. The team shortened the update to one sentence plus an offer to text the details, added a maximum call duration, and enabled voicemail detection for callbacks. Calls got shorter, callers were happier with the text link, and cost per resolved call dropped, without changing the model or voice at all.
5:28 Common mistakes
Common operations mistakes. Monitoring only volume and cost, not quality and compliance. Shipping changes to all traffic at once without a staged rollout or rollback. Optimizing cost per minute instead of cost per successful outcome, which can reward agents that hang up quickly without helping. And forgetting peak seasons, like Ramadan evenings, sale days or holiday periods, when concurrency limits and tool capacity really matter.
5:56 Watch me do it: week-one operations
Watch me do it. Aria has been live for a week. I open the analytics and set up our operating rhythm. First, alerts: any failed disclosure criterion alerts immediately; ninetieth percentile latency above budget alerts the owner; success rate down more than a set amount week on week alerts too. Next, the weekly review: I filter the fifty lowest-scoring conversations and read ten. Three patterns appear: callers asking about parking, which isn't in the knowledge base; a caller with a Pakistani accent misheard on the branch name; and long pauses when the booking API was slow on Monday morning. Each becomes an action: a parking FAQ chunk, a pronunciation and recognition hint for branch names plus a new audio test, and a cached availability fallback. Then cost: I fill the unit economics sheet with calls, average minutes, containment rate and current prices, and calculate cost per successful booking. I notice average call length is inflated by Aria reading the full cancellation policy, so I shorten it to one sentence with an offer to text the details. Finally, change management: the new prompt goes to twenty percent of calls through a branch with a traffic split, and I schedule the comparison for next Friday, with the rollback button noted in the runbook.
7:28 Recap and next step
Recap. After launch, monitor outcomes, friction, latency, tools, evaluation failures, per-language success and cost per outcome. Run a weekly feedback loop, version everything, roll out in stages and keep a rollback. Optimize cost per successful outcome with shorter turns, lean context and right-sized models, and plan capacity for peaks. Your next step: fill in the unit economics sheet in the lesson text with your own data and current vendor pricing, and set three alerts: disclosure failure, p90 latency and success-rate drop.
8:03 Try this now
Try this now. Open your platform's analytics and list the seven signals from this lesson. Set three alerts: any disclosure failure, ninetieth-percentile latency above budget, and a week-on-week drop in success rate. Fill in the unit economics sheet with your volumes and current vendor prices. Schedule a thirty-minute weekly review with a named owner, and create a place where every production failure becomes a new test case.
After launch, the real work starts
Real callers are more varied than any test suite. Production needs observability (what is happening), feedback loops (what to fix), cost control (unit economics) and safe change management (improving without breaking).
What to monitor
| Signal | Why | Alert when (illustrative) |
|---|---|---|
| Call volume and outcomes | Business health | Success rate drops vs last week |
| Transfer and abandonment rates | Friction | Abandonment spikes after a change |
| Voice-to-voice latency (median, p90) | Experience | p90 exceeds budget |
| Tool errors and timeouts | Integration health | Error rate above threshold |
| Evaluation criteria failures | Quality and compliance | Any disclosure or safety failure |
| Language mix and per-language success | Hidden gaps | One language falls below minimum |
| Cost per call and per successful outcome | Unit economics | Cost per outcome rises |
Feedback loops
- Weekly transcript review: sample calls, especially failures, transfers and low-score conversations.
- Topic and intent analysis: what are people calling about that the agent can't handle? (Some platforms cluster topics automatically.)
- Turn failures into tests: every production failure becomes a new test case.
- Knowledge gaps: unanswered questions become FAQ entries with owners.
Change management
- Version everything: prompt, tools, knowledge, voice, model and turn settings.
- Staged rollouts: route a percentage of traffic to a new version (some platforms support branches with traffic splits), compare outcomes, then promote.
- Rollback plan: one click back to the previous version.
- Model deprecations: providers retire models; track notices and test replacements early.
Understanding cost
Cost per minute for a cascaded agent is roughly:
cost/min ~= ASR + LLM tokens (input + output, per turn, grows with history and retrieved context)
+ TTS characters spoken + telephony (per-minute carrier/CPaaS rates, number rental)
+ platform fee (if bundled) + tool/API costsSpeech-to-speech models bill audio input and output tokens instead of separate ASR and TTS. Prices change often and differ by plan and region; always use current pricing pages. What matters is cost per successful outcome (per booking, per qualified lead, per resolved ticket), compared with the human alternative and the value of the outcome.
Levers to reduce cost without hurting quality
- Shorter agent turns: fewer TTS characters and output tokens; often better UX too.
- Lean context: trim conversation history, retrieve fewer chunks, use prompt caching where supported.
- Right-sized models: a fast, cheaper LLM for routine turns; escalate to a stronger model only when needed (routing or workflow nodes).
- Deflect cleanly: end calls politely when complete; enforce silence timeouts and maximum durations.
- Voicemail detection: do not pay for long messages to machines.
- Telephony choices: regional carriers and SIP trunking can change per-minute rates.
Scaling
- Concurrency limits: platforms and plans limit simultaneous calls; plan for peaks (Ramadan evenings for Gulf retailers, sale days, campaign launches).
- Tool capacity: your booking API must handle peak concurrency; load test it.
- Human backup: transfers spike when something breaks; ensure staff coverage or callback queues.
- Multi-region: for latency and data residency, deploy agents and tools near callers.
Worked example: a KSA e-commerce order-status line during a sale
Before the sale: load-tested the order API at triple normal concurrency; raised concurrency limits; added cached "sale FAQ" knowledge (delivery times, return policy); set max call duration; added an SMS link for detailed tracking to shorten calls. During: dashboard showed p90 latency creeping up as the order API slowed; the team enabled pre-tool speech and a cached status fallback. After: turned the ten most common failures into tests and new FAQ entries. Cost per resolved call fell because calls were shorter and containment rose.
Hands-on: a unit economics sheet
metric,value,notes
calls_per_month,12000,from platform analytics
avg_minutes_per_call,2.4,
all_in_cost_per_minute,<from current pricing>,ASR+LLM+TTS+telephony+platform
containment_rate,0.68,resolved without human
cost_per_call,=avg_minutes_per_call*all_in_cost_per_minute,
cost_per_resolved_call,=cost_per_call/containment_rate,
human_cost_per_call,<your data>,fully loaded
value_per_resolved_call,<your data>,e.g. saved agent time or revenueFill in from your own data and current vendor pricing; do not rely on anyone else's numbers.
Pitfalls
- Monitoring only volume and cost, not quality and compliance.
- Shipping changes to 100% of traffic without a staged rollout.
- Optimizing cost per minute instead of cost per successful outcome.
Key takeaways
- Monitor outcomes, transfers, abandonment, latency, tool errors, evaluation failures, per-language success and cost per outcome.
- Run weekly transcript reviews; turn failures into tests and knowledge gaps into FAQ entries.
- Version everything, use staged rollouts with traffic splits and keep one-click rollback; track model deprecations.
- Optimize cost per successful outcome with short turns, lean context, right-sized models, timeouts and voicemail detection; plan capacity for peaks.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Fill in the unit economics sheet with your data and current vendor pricing, and configure three alerts: any disclosure failure, p90 latency over budget, and a week-on-week success drop.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.