Voice AI & Conversational AgentsQuality, evaluation and operations · Lesson 14 of 17

Monitoring, cost control and scaling

Article · 15 min · 8 min lecture

Video lecture

Monitoring, cost control and scaling

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

After launch

  • Monitor
  • Feedback loops
  • Safe changes
  • Cost and scaling

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

After launch, the real work starts

Real callers are more varied than any test suite. Production needs observability (what is happening), feedback loops (what to fix), cost control (unit economics) and safe change management (improving without breaking).

What to monitor

SignalWhyAlert when (illustrative)
Call volume and outcomesBusiness healthSuccess rate drops vs last week
Transfer and abandonment ratesFrictionAbandonment spikes after a change
Voice-to-voice latency (median, p90)Experiencep90 exceeds budget
Tool errors and timeoutsIntegration healthError rate above threshold
Evaluation criteria failuresQuality and complianceAny disclosure or safety failure
Language mix and per-language successHidden gapsOne language falls below minimum
Cost per call and per successful outcomeUnit economicsCost per outcome rises

Feedback loops

  • Weekly transcript review: sample calls, especially failures, transfers and low-score conversations.
  • Topic and intent analysis: what are people calling about that the agent can't handle? (Some platforms cluster topics automatically.)
  • Turn failures into tests: every production failure becomes a new test case.
  • Knowledge gaps: unanswered questions become FAQ entries with owners.

Change management

  • Version everything: prompt, tools, knowledge, voice, model and turn settings.
  • Staged rollouts: route a percentage of traffic to a new version (some platforms support branches with traffic splits), compare outcomes, then promote.
  • Rollback plan: one click back to the previous version.
  • Model deprecations: providers retire models; track notices and test replacements early.

Understanding cost

Cost per minute for a cascaded agent is roughly:

cost/min ~= ASR + LLM tokens (input + output, per turn, grows with history and retrieved context)
          + TTS characters spoken + telephony (per-minute carrier/CPaaS rates, number rental)
          + platform fee (if bundled) + tool/API costs

Speech-to-speech models bill audio input and output tokens instead of separate ASR and TTS. Prices change often and differ by plan and region; always use current pricing pages. What matters is cost per successful outcome (per booking, per qualified lead, per resolved ticket), compared with the human alternative and the value of the outcome.

Levers to reduce cost without hurting quality

  1. Shorter agent turns: fewer TTS characters and output tokens; often better UX too.
  2. Lean context: trim conversation history, retrieve fewer chunks, use prompt caching where supported.
  3. Right-sized models: a fast, cheaper LLM for routine turns; escalate to a stronger model only when needed (routing or workflow nodes).
  4. Deflect cleanly: end calls politely when complete; enforce silence timeouts and maximum durations.
  5. Voicemail detection: do not pay for long messages to machines.
  6. Telephony choices: regional carriers and SIP trunking can change per-minute rates.

Scaling

  • Concurrency limits: platforms and plans limit simultaneous calls; plan for peaks (Ramadan evenings for Gulf retailers, sale days, campaign launches).
  • Tool capacity: your booking API must handle peak concurrency; load test it.
  • Human backup: transfers spike when something breaks; ensure staff coverage or callback queues.
  • Multi-region: for latency and data residency, deploy agents and tools near callers.

Worked example: a KSA e-commerce order-status line during a sale

Before the sale: load-tested the order API at triple normal concurrency; raised concurrency limits; added cached "sale FAQ" knowledge (delivery times, return policy); set max call duration; added an SMS link for detailed tracking to shorten calls. During: dashboard showed p90 latency creeping up as the order API slowed; the team enabled pre-tool speech and a cached status fallback. After: turned the ten most common failures into tests and new FAQ entries. Cost per resolved call fell because calls were shorter and containment rose.

Hands-on: a unit economics sheet

metric,value,notes
calls_per_month,12000,from platform analytics
avg_minutes_per_call,2.4,
all_in_cost_per_minute,<from current pricing>,ASR+LLM+TTS+telephony+platform
containment_rate,0.68,resolved without human
cost_per_call,=avg_minutes_per_call*all_in_cost_per_minute,
cost_per_resolved_call,=cost_per_call/containment_rate,
human_cost_per_call,<your data>,fully loaded
value_per_resolved_call,<your data>,e.g. saved agent time or revenue

Fill in from your own data and current vendor pricing; do not rely on anyone else's numbers.

Pitfalls

  • Monitoring only volume and cost, not quality and compliance.
  • Shipping changes to 100% of traffic without a staged rollout.
  • Optimizing cost per minute instead of cost per successful outcome.

Key takeaways

  • Monitor outcomes, transfers, abandonment, latency, tool errors, evaluation failures, per-language success and cost per outcome.
  • Run weekly transcript reviews; turn failures into tests and knowledge gaps into FAQ entries.
  • Version everything, use staged rollouts with traffic splits and keep one-click rollback; track model deprecations.
  • Optimize cost per successful outcome with short turns, lean context, right-sized models, timeouts and voicemail detection; plan capacity for peaks.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Which metric best reflects the economics of a booking agent?
  2. You want to test a new prompt safely in production. What should you do?
  3. A production call reveals a new failure pattern. What is the best long-term response?

Put it into practice

Fill in the unit economics sheet with your data and current vendor pricing, and configure three alerts: any disclosure failure, p90 latency over budget, and a week-on-week success drop.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.