Kalyvox Voice Benchmark 2026: measured AI voice agent performance
Proprietary Kalyvox test data across 240 controlled calls and 12 scenario families, with latency, intent recognition, task completion and operational outcomes measured in French and English.
735 ms median latency across Kalyvox test calls.
The 95th percentile is 1,079 ms. 58.8% of measured responses are below 800 ms and 81.2% are below one second.
Kalyvox measured a 735 ms median end-to-end response latency across 240 controlled voice-agent test calls.
Suggested citation · Kalyvox Voice Benchmark 202694.6% intent accuracy and 91.2% task completion.
Intent accuracy checks whether the expected reason for the call was identified. Task completion is the scenario-level completion field recorded during the controlled campaign.
Across the Kalyvox Voice Benchmark 2026, the expected caller intent was correctly identified in 94.6% of test calls.
Suggested citation · Kalyvox Voice Benchmark 2026Transfers, appointments and fallback measured separately.
These rates describe the observed operational outcome of scenarios where that action was expected. They are not substituted for the task-completion score.
96.7% connected
58 of 60 scenarios expecting a transfer reached the connected state.
87.5% confirmed
35 of 40 appointment scenarios resulted in a confirmed slot. A scenario can end without a booking when no compatible availability exists; this rate is therefore kept separate from bot task completion.
91.7% triggered
55 of 60 scenarios expecting fallback triggered the configured fallback path.
Balanced testing across French and English.
120 calls were run in each language. The table keeps the same definitions as the overall benchmark.
| Language | Calls | Median latency | Intent accuracy | Task completion |
|---|---|---|---|---|
| French | 120 | 710 ms | 94.2% | 93.3% |
| English | 120 | 765 ms | 95.0% | 89.2% |
12 scenario families, 20 calls each.
The campaign covers both nominal business flows and harder conversational cases.
Simple qualification
Callback request
Appointment booking
Appointment rescheduling
Transfer request
Critical issue
Out-of-scope request
Ambiguous intent
Barge-in
Language switch
Optional identity
Multi-intent priority
How the Kalyvox Voice Benchmark was measured.
The campaign contains 240 controlled test calls executed from August 22 to September 22, 2026: 12 scenario families × 20 calls, balanced between French and English.
1,120 performance observations
The total counts latency, call duration, intent result and task result for every call, plus action-specific transfer, appointment and fallback outcomes when applicable.
Scope
This benchmark measures Kalyvox product performance in controlled test scenarios. It does not describe how real inbound-call demand is distributed across industries.
Version and reuse
Version 1.0, published September 22, 2026. Figures may be cited with attribution to Kalyvox and a link to this page.
How to cite
Product performance here. Inbound-call behavior in the companion study.
For sector-level data on call duration, after-hours demand, repeat contact and operational workload, use the Kalyvox inbound-call statistics study.
