NVIDIA Releases Open-Weight Nemotron 3 Diarization Model
NVIDIA released a 100M-parameter open-weight Nemotron 3 Diarization model that tracks up to eight overlapping speakers in real time.
On September 23, 2026, NVIDIA released Nemotron 3 Diarization, a real-time model for identifying who spoke when in multi-speaker audio streams. MarkTechPost reports it is a 100-million-parameter open-weight model on Hugging Face under the OpenMDW License 1.1, which permits commercial use, with one checkpoint covering offline recordings and streaming latency down to 0.32 seconds. That checkpoint tracks up to eight speakers, including overlapping speech, doubling the four-speaker limit of Streaming Sortformer. On Voice Arena's initial Diarization-Bench of 139 English conversations totaling about 22 hours, it ranked first among 12 systems at 14.72% DER versus 19.3% for the next system. It runs through NVIDIA NeMo on Ampere, Ada Lovelace, Hopper, or Blackwell GPUs, and MarkTechPost says more than eight speakers, heavy noise, or far-field audio can raise error rates. The Hugging Face blog describes the same release more generally and does not dispute those figures.
- On September 23, 2026, NVIDIA released Nemotron 3 Diarization, a real-time model for identifying who spoke when in multi-speaker audio.
- It is a 100-million-parameter open-weight model on Hugging Face under the OpenMDW License 1.1, which permits commercial use.
- One checkpoint tracks up to eight overlapping speakers for offline recordings and streaming down to 0.32 seconds, doubling Streaming Sortformer's four-speaker limit.
- On Voice Arena's initial Diarization-Bench of 139 English conversations totaling about 22 hours, it ranked first among 12 systems at 14.72% DER versus 19.3% for the next system.
- It runs through NVIDIA NeMo on Ampere, Ada Lovelace, Hopper, or Blackwell GPUs.
- More than eight speakers, heavy noise, or far-field audio can raise error rates.
Coverage timelineoldest first · each row is one article
- · 3d ago**Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**
Hugging Face Blog· 50
NVIDIA released Nemotron 3 Diarization, a real-time AI model for identifying multiple speakers in audio streams.
- · 3d agoNVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
MarkTechPost· 54
NVIDIA released Nemotron 3 Diarization, a 100M open-weight model tracking up to eight overlapping speakers.