The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Model Releases · Nvidia

Nvidia releases free model for real‑time speaker ID

Nemotron 3 Diarization has open weights on Hugging Face and can tell apart up to eight speakers live, which Nvidia says cuts errors by 41 percent against its predecessor.

Nvidia released Nemotron 3 Diarization, a 100-million-parameter model that identifies who is speaking among up to eight people in real time. The weights are free on Hugging Face under an open license, The Decoder reported.

The model topped VoiceArena's Diarization-Bench v1 with a 14.72 percent error rate. Nvidia's model card puts the gain over its prior Streaming Sortformer model at an average 41 percent across eight test scenarios, using a roughly one-second buffer.

On the DIHARD III benchmark, Nvidia said the model scored a 12.73 percent error rate against 19.09 percent for its baseline. Streaming inference runs as fast as 38 times real-time at a 1.04-second buffer, or 12.5 times real-time at 0.32 seconds, the model card says.

Nvidia said accuracy falls off past eight speakers, in heavy background noise or with room echo. Paired with the company's Parakeet speech-recognition model, the tool can label a transcript by speaker without naming who anyone actually is.

Sources 2 sources

  1. Source The Decoder
  2. Source Nvidia (Hugging Face model card)