StepFun speech model tops Artificial Analysis transcription index at 1.7%
The evaluator said StepAudio 3 ASR cuts word error rate from 4.7% to 1.7%, but costs more per minute than four rivals.
StepFun's StepAudio 3 ASR ranks first on the AA-WER Index for non-streaming speech to text at a 1.7% word error rate, Artificial Analysis said on Tuesday. The evaluator put the previous model, StepAudio 2.5 ASR, at 4.7% on the same index.
The score is the first time a StepFun model has reached the top of that leaderboard, according to Artificial Analysis. The model is available through the StepFun API for non-streaming transcription. The figures are the evaluator's own measurements, run on its index rather than supplied by StepFun.
Price runs the other way. StepAudio 3 ASR costs $0.40 per hour of audio, or $6.67 per 1,000 minutes, Artificial Analysis said. The evaluator called that the most expensive of the five most accurate models on its non-streaming leaderboard.
MAI-Transcribe-2 and Grok Voice Transcribe 2.0 cost $1.67 per 1,000 minutes on the evaluator's figures, and ElevenLabs Scribe v2 costs $3.67. On that account the model trades a higher price per minute for its accuracy lead on conversational audio.
On speed the model sits mid-table. Artificial Analysis put StepAudio 3 ASR at 88 times real time. That is behind MAI-Transcribe-2 at 374 times and Grok Voice Transcribe 2.0 at 154 times, and roughly level with ElevenLabs Scribe v2 at 84 times.