Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real time
Nvidia released free 100-million-parameter Nemotron 3 Diarization, which labels up to eight speakers in real time.
Nvidia released Nemotron 3 Diarization, an approximately 100-million-parameter model with freely available weights that labels who is speaking. It handles up to eight speakers, including overlap, on recordings or live audio, and can pair with Parakeet for anonymously labeled transcripts. On VoiceArena Diarization Benchmark v1 it reports a 14.7% diarization error rate, about 41% lower than Streaming Sortformer at a 1.04-second buffer, ahead of the next system at 19.3%. Buffers can be set from 30.4 seconds down to 0.32 seconds, and shorter buffers generally reduce accuracy.