Speech AI
Voice AI said 'I understand your frustration.' It had no idea what that meant.
Hume's Real World VoiceEQ benchmark, based on more than one million human ratings, tests over 40 voice models across dimensions standard benchmarks ignore: emotion, speaker identity, and acoustic context. The findings show speech-to-speech models vary wildly, and even leading systems often ignore the audio cues humans use instinctively.
Emmanuel Fabrice Omgbwa Yasse AI-assisted
2026-07-20 · 4 min read

Voice is supposed to be AI's next big interface. Customer support, healthcare, education, personal assistants, the pitch has been consistent for years. On paper, things look good. Word error rates have dropped. Latency hits conversational speeds. The usual benchmarks are close to saturated.
Then you talk to one. The model that starts the conversation sounding like one person might end it sounding like someone else. The banking agent that cheerfully says "I understand your frustration" with no trace of empathy. The pause the system fills with a generic answer because it missed the hesitation entirely. This is the gap standard benchmarks miss, and it's the one users actually feel. A similar blind spot affects AI coding assistants, where metrics like acceptance rate don't capture whether a suggestion is actually correct, as one analysis of Devin's human-hour savings found.
Hume, the company behind the empathy-focused voice AI platform, built a new evaluation suite called Real World VoiceEQ designed to measure exactly what those metrics miss. The results paint a messy picture of a field that optimized speed and accuracy at the expense of something harder to score: whether the voice on the other end actually sounds human.
The method: one million ratings, 40+ models
The benchmark evaluates more than 40 proprietary and open-source voice models across 15 evaluation dimensions and more than 60 metrics, covering ASR, TTS, speech-to-speech, and speech understanding. The data comes from over one million human ratings collected across different demographics, speaking styles, and acoustic environments. The current release includes 785,000 TTS ratings and 48,000 STS ratings, among the largest human evaluations of voice AI conducted publicly. For comparison, even leading automated evaluation systems struggle with subjective voice quality, echoing the broader challenge that AI agents hit 49% on a test human experts pass at 95%.

Ratings were collected using Hume's evaluation platform, Kairos, which frontier labs also use to run custom evaluations and generate human preference data.
No single winner
The headline finding is one practitioners have suspected for a while: there is no best voice model. In TTS evaluation, no system configuration ranked among the top five across all eight capability groups. One model nails pronunciation of complex pharmaceutical names and bank account numbers, but produces wooden, monotone speech. Another sounds remarkably natural but flakes on precision tasks.
The variation is most extreme in speech-to-speech models, which handle recognition and generation in a single pipeline. Some recognize emotion well but cannot respond naturally. Others stay largely transcript-driven, relying on the words spoken while ignoring tone, pacing, hesitation, emphasis, and volume. This fragmented landscape reminds me of the evaluation problem in AI agents more broadly, where teaching agents to be better judges of their own outputs has become the bottleneck.
That last point is the one that stings. Humans use those cues constantly. A confident "yes" and a hesitant "...yes..." carry completely different meanings, but the transcript looks identical. The benchmark found that even access to audio did not guarantee agents used the paralinguistic information it contained.
Where the numbers hide real problems
The benchmark also shows how aggregate scores mask failure modes. In one example, transcription word error rates on noise-backed speech were roughly four times higher than on music-backed speech. A single background-audio score would have flattened that difference and made the failure invisible. Models optimized for public benchmarks can overfit to clean conditions, a problem that also surfaces in other domains, for instance, one CIFAR-10 fix did nothing for MNIST, challenging how researchers interpret benchmark results.
In preliminary research, Hume also found signs that models may be optimized for public benchmarks rather than real conditions. Several reproduced known errors in reference transcripts, followed arbitrary spelling conventions, and even reconstructed masked words not present in the audio.
Automated evaluation isn't ready yet
LLMs are widely used to evaluate text models. Hume's findings suggest that speech-language models (SLMs) should be used cautiously for voice evaluation. When the company compared leading SLMs with trained human raters on TTS assessments, agreement was highest on tasks with clear, verifiable answers, such as pronunciation accuracy. Agreement declined sharply on more subjective evaluations. SLMs sometimes appeared to infer emotion from text-based contextual cues rather than acoustic ones. The weakest agreement came from open-ended judgments like whether a voice fit an acting role or maintained consistent identity over a conversation. This echoes the finding that faster LLM training kernels don't automatically solve evaluation reliability.
Automated evaluators are useful for well-defined tasks. They are not a substitute for human listeners when context, perception, and social interpretation are on the line.
Hume positions Real World VoiceEQ as an extension of the traditional evaluation paradigm: a human-grounded metric for the components of synthetic voice that quantitative scores have never captured well. The full technical report and leaderboards are public. The company also offers custom evaluations through its Kairos platform.
- Source : Voice AI is getting better at words, but it still can't hear how you say them — 2026-07-15
Get the tech essentials in 3 minutes every morning
One email, every weekday, with what actually matters in AI and tech.