返回全部动态

AWS 在 SageMaker AI 上推出 WhisperX 说话人标注转写容器

原标题:Speaker-labeled transcription with WhisperX on SageMaker AI

AWS Machine Learning Blog一手来源产品发布质量 69

AI 摘要

AWS 发布 WhisperX 深度学习容器(DLC),在 SageMaker AI 上提供带说话人标注的语音转写能力。该容器封装了 OpenAI Whisper,并加入 wav2vec2 强制对齐实现逐词时间戳,以及说话人分离(diarization)功能,无需 Hugging Face token 即可部署。用户可将其部署到 SageMaker AI 实时或异步端点,实时端点适合 60 秒内短音频,异步端点适合长音频和高吞吐批处理,并支持 S3 输入输出与自动扩缩容。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Speaker-labeled transcription with WhisperX on SageMaker AI Any team working with spoken audio hits the same wall with generic speech-to-text. Think contact-center calls, all-hands meetings, podcasts, depositions, and broadcast media. These workloads need two things that standard transcription gets wrong. First, timestamps land at the utterance level, off by several seconds. Second, there’s no reliable answer to “who said what.” Those gaps make transcripts hard to search, caption, redact, or ana


发布时间:2026-09-25 00:20
抓取时间:2026-09-25 01:02
来源机构:AWS
阅读原文aws.amazon.com