The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Model Releases · Artificial Analysis · Cartesia · ElevenLabs · Mistral AI

Artificial Analysis opens speech leaderboards for 9 languages beyond English

The evaluator said no single model leads everywhere, with Cartesia's Sonic 3.6 top in 4 languages and four labs splitting the open-weights lead.

Artificial Analysis opened Multilingual Text to Speech Arena leaderboards on Tuesday, covering nine languages beyond English. The evaluator said performance varies across languages, and that models which do well in English do not necessarily deliver the same pronunciation and naturalness elsewhere.

Cartesia's Sonic 3.6 leads Hindi and Arabic, with a 33-point Elo lead in Hindi and a 132-point lead in Arabic over the next model, according to Artificial Analysis. ElevenLabs' Eleven v3 ranks second in both, and v3 Conversational third in both.

In Hindi the evaluator put Sonic 3.6 at 1,179 Elo and Eleven v3 at 1,146. Sonic 3.6 also leads Japanese and Vietnamese. Inworld's Realtime TTS-2 ranks first in Mandarin, where StepFun's StepAudio 2.5 TTS is third at 1,130 Elo, 16 points behind Sonic 3.6.

No single open-weights model leads across the set. Boson AI's Higgs Audio V3 TTS is top on open weights in four of the nine languages, among them Hindi and Portuguese, the evaluator said. FishAudio's OpenAudio S1 Mini leads on open weights in Japanese and Arabic, and Mistral's Voxtral TTS leads the rest.

The rankings are Artificial Analysis's own. The Elo scores come from its arena rather than from the labs, and it published them in the same week as a separate English pronunciation test. The evaluator has not said how many votes each language's leaderboard rests on.

Sources 4 sources

  1. Source ArtificialAnlys
  2. Source ArtificialAnlys
  3. Source ArtificialAnlys
  4. Source ArtificialAnlys