The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Model Releases · Alibaba

Alibaba releases 5 Qwen audio models, cuts speech recognition price

Qwen said the upgrade covers recognition, synthesis and realtime conversation, adds two models, and cuts prices by up to 95%.

Alibaba's Qwen team released Qwen-Audio-3.1 on Wednesday, an upgrade covering speech recognition, text-to-speech and realtime conversation, with two models added. Qwen said prices fall across the line: text-to-speech by about 70%, realtime by about 85%, and speech recognition by up to 95%.

The pricing is the company's own. Qwen Cloud lists the ASR Flash file-transcription model at $0.15 per million input tokens and $0.47 per million output tokens, with a rate limit of 600 requests a minute. The announcement gives the percentage cuts without the prior prices behind them.

The two new models are ASR-Next and TTS-Next. Qwen said ASR-Next handles multi-speaker recognition with speaker labels and timestamps, and can caption ambient and machine sounds. It said TTS-Next combines a language model with diffusion to generate voice, sound effects and background audio in a single pass.

Qwen said the existing recognition model gained multilingual and dialect coverage and a polishing step that strips fillers and repetitions from transcripts, and that a user can interrupt the realtime model mid-turn. No independent benchmark has tested any of these claims in this material.

Qwen also said the realtime model slows down and responds empathetically when it senses a low mood. The announcement does not say what triggers that behaviour or how the company measures it. The company said more APIs are coming, and pointed to Qwen Cloud listings for the recognition and realtime models.

Sources 2 sources

  1. Source Alibaba_Qwen
  2. Source Qwen Cloud