返回全部动态

NVIDIA 发布 Nemotron 3 Diarization 实时多说话人模型

原标题:**Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization**

Hugging Face Blog一手来源模型发布质量 73

AI 摘要

NVIDIA 发布开源权重模型 Nemotron 3 Diarization,参数量 1 亿,在 VoiceArena Diarization-Bench 排行榜以 14.72% 的 DER 位列第一。该模型支持最多 8 位说话人、重叠语音、分块处理和可定制流式延迟,并采用到达顺序排序与 AOSC、FIFO 记忆机制保持说话人标签稳定。它可与 ASR 结合生成带说话人归属的转录,提升会议、客服和语音代理等场景的可用性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Turn overlapping conversations into speaker-aware data with one open-weight, 100M-parameter model - ranked #1 in VoiceArena's Diarization leaderboard with a 14.72% Diarization Error Rate (DER). Every conversation carries two layers of information: what was said and who said it. Speech recognition captures and transcribes the words. Speaker diarization classifies who spoke when, helping applications connect what was said to the right participant. Consider a transcript from a meeting, customer cal


发布时间:—
抓取时间:2026-09-24 02:52
来源机构:Hugging Face
阅读原文huggingface.co