NVIDIA 发布 Nemotron 3 Diarization 实时多说话人模型
原标题:**Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**
AI 摘要
NVIDIA 发布开源权重模型 Nemotron 3 Diarization,参数量 1 亿,在 VoiceArena Diarization-Bench 排行榜以 14.72% 的 DER 位列第一。该模型支持最多 8 位说话人、重叠语音、分块处理和可定制流式延迟,并采用到达顺序排序与 AOSC、FIFO 记忆机制保持说话人标签稳定。它可与 ASR 结合生成带说话人归属的转录,提升会议、客服和语音代理等场景的可用性。
正文节选
Turn overlapping conversations into speaker-aware data with one open-weight, 100M-parameter model - ranked #1 in VoiceArena's Diarization leaderboard with a 14.72% Diarization Error Rate (DER). Every conversation carries two layers of information: what was said and who said it. Speech recognition captures and transcribes the words. Speaker diarization classifies who spoke when, helping applications connect what was said to the right participant. Consider a transcript from a meeting, customer cal