Nvidia drops a free 100M-parameter model that identifies up to eight speakers in real time
Nvidia released free 100-million-parameter Nemotron 3 Diarization, which labels up to eight speakers in real time.
Nvidia released Nemotron 3 Diarization, an approximately 100-million-parameter model with freely available weights that labels who is speaking. It handles up to eight speakers, including overlap, on recordings or live audio, and can pair with Parakeet for anonymously labeled transcripts. On VoiceArena Diarization Benchmark v1 it reports a 14.7% diarization error rate, about 41% lower than Streaming Sortformer at a 1.04-second buffer, ahead of the next system at 19.3%. Buffers can be set from 30.4 seconds down to 0.32 seconds, and shorter buffers generally reduce accuracy.
- About 100 million parameters, with weights released for free.
- Separates up to eight speakers and detects overlapping speech.
- VoiceArena DER is 14.7%, roughly 41% below Streaming Sortformer.
- Buffers span 0.32 to 30.4 seconds; shorter windows hurt accuracy.
- With Parakeet it can add anonymous speaker labels to transcripts.
Full article262 words · extracted from the-decoder.com · click to collapse
Sep 27, 2026
Nvidia released Nemotron 3 Diarization, an AI model that identifies which speaker is talking at any given moment in a conversation. The model has about 100 million parameters, and its weights are freely available. It can tell apart up to eight speakers and detect when multiple people talk at the same time. More participants, heavy background noise, or reverb push error rates higher. Paired with a speech recognition system like Parakeet, the model can produce transcripts with speaker labels, though only anonymous ones like "speaker_2." It works with both recordings and live audio.

The audio buffer can be set to four levels ranging from 30.4 down to 0.32 seconds. Shorter buffers generally reduce accuracy. On the Diarization-Bench from VoiceArena, the model currently sits in first place with a 14.72 percent error rate, ahead of the next best system at 19.3 percent. The benchmark is strict. Overlapping speech counts, and even tiny misalignments at speaker transitions are scored as errors. Compared to its predecessor, Streaming Sortformer, the new model cuts the error rate by an average of 41 percent across eight test scenarios when using a 1.04-second buffer.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Text extracted automatically; images, tables and formatting may be missing. Original: