nvidia/Nemotron-3-Diarization

NVIDIA releases Nemotron 3 Diarization, an open-weight model identifying up to eight speakers in audio. It supports streaming inference with 80 ms latency and offline processing. The model uses Sortformer architecture with Arrival-Order Speaker Cache. Developers can install NVIDIA NeMo Speech to run inference on WAV files. The release includes a live demo and commercial usage rights.

NVIDIA released Nemotron 3 Diarization, an open-weight model that identifies up to eight speakers in audio streams. The system determines who spoke when by resolving speaker permutations based on the order of first arrival. It utilizes the Sortformer architecture, which includes a speaker cache to maintain identity over time. Developers install NVIDIA NeMo Speech to run inference on WAV files or Python numpy arrays. The release supports both streaming processing with 80 ms buffer latency and offline analysis with a 30.4 s buffer. Input resolution is configurable in multiples of 10 ms. The September 23, 2026 release includes a live demonstration page. NVIDIA states the model is ready for commercial use, but users must verify local hardware compatibility. The 80 ms latency figure excludes computational processing time, so total response delays may be higher. Users should confirm that the specific streaming configuration matches their application requirements. The card notes that Python 3.12 or later is required for installation.

README

nvidia/Nemotron-3-Diarization View on Hugging Face

Loading the README from Hugging Face…