AWS 在 SageMaker AI 上推出 WhisperX 说话人标注转写容器
原标题:Speaker-labeled transcription with WhisperX on SageMaker AI
AI 摘要
AWS 发布 WhisperX 深度学习容器(DLC),在 SageMaker AI 上提供带说话人标注的语音转写能力。该容器封装了 OpenAI Whisper,并加入 wav2vec2 强制对齐实现逐词时间戳,以及说话人分离(diarization)功能,无需 Hugging Face token 即可部署。用户可将其部署到 SageMaker AI 实时或异步端点,实时端点适合 60 秒内短音频,异步端点适合长音频和高吞吐批处理,并支持 S3 输入输出与自动扩缩容。
正文节选
Speaker-labeled transcription with WhisperX on SageMaker AI Any team working with spoken audio hits the same wall with generic speech-to-text. Think contact-center calls, all-hands meetings, podcasts, depositions, and broadcast media. These workloads need two things that standard transcription gets wrong. First, timestamps land at the utterance level, off by several seconds. Second, there’s no reliable answer to “who said what.” Those gaps make transcripts hard to search, caption, redact, or ana