ElevenLabs' Eleven v4 tops independent voice AI benchmark
Artificial Analysis says ElevenLabs' new Eleven v4 text-to-speech model ranks first on its Provider Voice leaderboard, beating Cartesia and Google models on pronunciation and speed.
Artificial Analysis said ElevenLabs' new Eleven v4 text-to-speech model ranks first on its Provider Voice TTS Arena leaderboard, beating rival models from Cartesia and Google.
https://x.com/ArtificialAnlys/status/2104578736687653293
The model scored an Elo of 1,319 across 1,674 head-to-head appearances, the benchmark firm said. That put it ahead of Cartesia's Sonic 3.6, at 1,276, and Google's Gemini 3.8 Flash TTS, at 1,267. Eleven v4 also topped all four Provider Voice categories: customer service, assistants, knowledge sharing and entertainment.
On the benchmark measuring how well the model pronounces difficult words, Eleven v4 scored 91.7%, the highest Artificial Analysis said it has measured, up from 85.6% for the prior Eleven v3. The new model also processes speech at 73.4 characters per second, nearly double Eleven v3's rate. Though speed improved, Eleven v4 also costs more: $80 per million characters, against $49 for Sonic 3.6 and $16.49 for Gemini 3.8 Flash TTS.
On a controlled-voice test, where every model narrates with the same custom voice, Eleven v4 ranked second behind Alibaba's Qwen-Audio-3.1-TTS-Plus, Artificial Analysis said. The firm said Eleven v4 also expands language support to more than 90 languages, up from 70-plus for Eleven v3. ElevenLabs itself has not published its own account of the release.