返回全部动态

Cue2Narrate:重新思考电影音频描述生成

原标题:From Visual Cues to Spoken Narration: Rethinking Audio Description

arXiv cs.CV一手来源研究质量 79

AI 摘要

该研究提出Cue2Narrate,一种两阶段音频描述生成管道,用于长电影片段,同时预测视觉线索窗口和语音叙述窗口。作者引入LongLSMDC基准,包含长达8分钟的电影片段,并证明Cue2Narrate在定位和生成性能上优于基线。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

From Visual Cues to Spoken Narration: Rethinking Audio Description Abstract Audio Description (AD) provides spoken narration of visual events during dialogue gaps, making movies accessible to visually impaired audiences. The problem requires determining both what (which visual event) and when (position for inserting the AD) to narrate, to achieve the best user experience. Prior work has largely reduced the problem to video captioning of pre-segmented video clips, i.e., what is largely predefined


发布时间:2026-09-03 12:00
抓取时间:2026-09-03 12:12
来源机构:arXiv
阅读原文arxiv.org