speech-to-speech
2 published articles
Benchmarks & Tests5 min read
Voice Agents
Why Grok 4.1 Fast beats smarter models on a phone call
Overall leaderboards pick the wrong LLMs for voice calls. BenchLM's 2026 ranking puts latency first: Grok 4.1 Fast answers in 0.54s while Claude Opus 4.6, the top scorer, needs 1.78s. Fast models take the conversation; reasoning models stay on background tool calls.
2026-08-05
AIFeatured4 min read
Speech AI
Voice AI said 'I understand your frustration.' It had no idea what that meant.
Hume's Real World VoiceEQ benchmark, based on more than one million human ratings, tests over 40 voice models across dimensions standard benchmarks ignore: emotion, speaker identity, and acoustic context. The findings show speech-to-speech models vary wildly, and even leading systems often ignore the audio cues humans use instinctively.
2026-07-20