nvidia/nemotron-diarization
NVIDIA publishes a Hugging Face Space demonstrating live microphone diarization and multilingual streaming ASR. The CPU-only environment supports audio files up to two minutes and includes conversation labs for stress testing. It uses Nemotron 3.5 for language detection and separate API routes for English and multilingual processing. The deployment relies on server-side credentials and deterministic launch scripts for scenario generation.
NVIDIA deploys Hugging Face Spaces demonstrating automatic speaker identification from live microphone input or uploaded recordings. The system utilizes the Nemotron 3.5 model for language detection and separate processing routes for multilingual and English audio. Audio files are limited to a duration of two minutes, while live microphone sessions automatically terminate after thirty seconds. A CPU-only environment handles these tasks without specialized GPU hardware on the client side. Individuals upload supported audio formats or activate microphone controls to begin transcription. The interface displays speaker activity lanes and transcripts in real-time as audio streams to the server. Users can select specific audio ranges for processing before transmission starts. Additional lab environments generate synthetic conversations for testing purposes using deterministic launch scripts. Developers must configure server-side credentials and specific environment variables for multilingual support. The multilingual tab remains disabled if its function identifiers are not properly set. Session limits apply to total concurrent uses within the deployment. NVIDIA states that external credentials stay in the runtime environment and are never transmitted to the browser.
README
nvidia/nemotron-diarization View on Hugging Face
Loading the README from Hugging Face…