How It Works Use Cases Features Pricing Blog About
Sign In Start Free Trial
Back to Blog

How We Test Voice Agents Before Going Live

Engineer reviewing a voice agent test transcript on a laptop

One of the things I learned building real-time audio pipelines is that the bugs that matter most are the ones that happen in the first few seconds. A system that works 95 percent of the time but fails unpredictably on the opening exchange is a system that callers will not trust, no matter how well it performs after that.

Voice agents have the same property. The first 10 seconds of a call set the caller's expectation for everything that follows. If the opening sounds robotic, cuts off mid-word, or takes a half-second too long to respond, the caller's guard goes up. They become less willing to give good answers to qualification questions, less forgiving of minor imperfections later in the flow, and more likely to ask for a human even if the agent could have handled their request fine.

Our pre-launch testing protocol is designed to catch those first-10-second problems before any real caller experiences them. Here is the full process.

Phase 1: Isolated script testing

Before we ever connect an agent to a real phone number, we run through the entire script in text-only mode. This is not as glamorous as live call testing, but it finds a category of bugs that audio testing often misses: logical gaps in the script itself.

We run three types of script tests:

  • Happy path: A caller who follows the most common flow from start to finish. Does the script have a natural resolution? Are there any points where the agent would be waiting for input that it cannot reasonably expect?
  • Short responses: A caller who responds with single words or very brief answers. "Yes." "No." "HVAC." The script needs to handle these without getting confused or looping back unnecessarily.
  • Divergent responses: A caller who answers a question with something tangential to what was asked. "Can I get your name?" / "Well, I've been a customer for three years." The agent should acknowledge and redirect, not get stuck.

Any script gaps found here get fixed before moving to audio testing. This saves significant time, because fixing a script-level problem after you have already tested audio means redoing the audio test anyway.

Phase 2: Voice quality and latency testing

Once the script is structurally clean, we test the audio layer. The two properties that matter most are voice naturalness and response latency.

Voice naturalness is subjective but testable. We have an internal rubric: play the opening 30 seconds of an agent call to someone who has not heard it before, without telling them what it is. If they immediately identify it as a robot, the voice configuration needs work. If they need a few exchanges to be sure, it is acceptable. If they are uncertain through the first minute, it is good.

The specific things we listen for: unnatural pronunciation of business-specific terms (your business name, your city, any service names you use), pacing that feels either rushed or artificially slow, and prosody on questions versus statements. A question that ends on a flat tone rather than a rising tone sounds like an automated system reading a script, even if the words are correct.

Response latency is more technical. We measure the gap between when the caller stops speaking and when the agent starts its response. Under 800 milliseconds feels natural. Over 1.2 seconds starts to feel like lag. Over 2 seconds causes callers to re-speak or ask "hello, are you there?" which breaks the conversational flow entirely. If our latency tests are coming back over 1 second consistently, we look at the inference pipeline before going live.

Phase 3: Live adversarial testing

This is the most valuable phase and the one that catches the most real bugs. We call the agent ourselves, but we do not cooperate. We specifically try to break it.

The adversarial test scenarios we run for every agent before launch:

  • Interruption test: Start speaking before the agent finishes its opening line. Does the agent handle barge-in correctly, or does it keep talking over you?
  • Silence test: Say nothing for 5 seconds after the agent asks a question. Does it reprompt appropriately, or does it loop endlessly or disconnect?
  • Profanity or anger test: Express frustration or use aggressive language. Does the agent escalate to a human appropriately, or does it try to continue the standard flow?
  • Wrong domain question: Ask something completely outside the script. "What's the weather like in Miami today?" or "Can you help me with my cat?" The agent should acknowledge it cannot help with that and redirect, not attempt an answer.
  • Transfer failure test: Deliberately trigger a transfer to a number that does not answer. Does the fallback chain activate correctly? Is the caller's information captured before the call ends?
  • Repeat caller test: Call in with the exact same information as a previous test call. Does the agent handle repeat contacts gracefully, or does the duplication cause issues in the CRM log?

Phase 4: Stakeholder call review

After adversarial testing, we run what we call a stakeholder review: we have someone from the business, typically the owner or whoever will be most affected by the agent's quality, call in and experience it as a first-time customer would. Not with adversarial intent, but with fresh ears.

The reason this step is distinct from internal testing is that business owners often notice things we do not. They know exactly how their company name should be pronounced. They know which phrases are specific to their local market and which sound generic. They know if the agent's opening line matches the tone they have been trying to build in their other customer communications.

We give stakeholders a structured feedback form: rate the opening, rate the naturalness, flag any specific phrases that felt wrong, note any points where they felt confused about what to do next. This structured collection process is more useful than open-ended "what did you think" feedback, because it catches specific fixable issues rather than general impressions.

The go/no-go criteria

We have a short list of blockers that prevent going live regardless of how well other tests went:

  • Response latency above 1.5 seconds on average across 10 test calls
  • Any test call that ends in an unresolved state without capturing contact information
  • Business name or key terms mispronounced in the opening
  • Any failure in the transfer fallback chain
  • Stakeholder feedback of "this would embarrass us if a real customer called"

If any of these are present, we fix and retest. We do not ship with known blockers. The cost of a bad first impression on a real customer is higher than the cost of another day of testing.

That said, we are not perfectionists. Minor imperfections in the middle of a flow, things a caller would barely notice, do not hold up a launch. The goal is a clean opening, a reliable resolution path, and a graceful exit for anything unexpected. Everything else can be refined once the agent is live and you have real call data to work with.

Put these ideas to work today

Start a free 14-day trial. Your voice agent goes live in under 5 minutes.