The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Model Releases · Artificial Analysis · Google · SpaceXAI

Artificial Analysis puts Gemini 3.1 Flash top of new speech benchmark

The evaluator's new pronunciation test measures how often text-to-speech models read difficult text correctly, and puts Google's entry at 88.1 percent.

Artificial Analysis said on Tuesday it has begun scoring text-to-speech models on what it calls "pronunciation robustness", a measure of how often a model reads challenging text correctly. The evaluator put Google's Gemini 3.1 Flash TTS top at 88.1 percent, at a listed $18.31 per million characters.

SpaceXAI TTS scored 87.6 percent at $15 per million characters, which Artificial Analysis called the lowest price of any model above 80 percent. Kokoro 82M v1.0 is the cheapest the evaluator listed, at $0.65 per million characters. Those are list prices, according to the evaluator, and not measured cost.

Speed and accuracy do not track each other in the results. Kokoro 82M v1.0 is the fastest model Artificial Analysis posted, at 242 characters a second. Realtime TTS-2 Flash reached 220 characters a second and scored 77.8 percent. SpaceXAI TTS ran at 106 characters a second.

The scores diverge from listener preference. Sonic 3.6 holds the highest Provider Voice Arena Elo of the group at 1276 while scoring 74.5 percent on the new test, the evaluator said. Gemini 3.1 Flash TTS sits at 1201 Elo. Both rankings are its own.

The test breaks into sub-scores. Artificial Analysis said Gemini 3.1 Flash TTS leads on contextually appropriate pronunciation at 96.4 percent and on expanding shorthand at 84.4 percent, against 84.3 percent for SpaceXAI TTS. SpaceXAI TTS leads on preserving exact sequences at 85.7 percent.

Sources 4 sources

  1. Source ArtificialAnlys
  2. Source ArtificialAnlys
  3. Source ArtificialAnlys
  4. Source ArtificialAnlys