StrAD:长视频音频描述生成的流式方法与基准
原标题:StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos
AI 摘要
StrAD 论文提出了一种用于长视频音频描述生成的新基准和流式方法,将任务重新定义为流式密集视频描述,无需真实时间戳即可在滑动窗口内生成 AD。其微调模型 StrAD-FT 在 CMD-AD 上达到 SOTA(CIDEr 36.3),在 StrAD 基准上为 51.0,但流式任务中时间定位和叙事连贯性仍有局限。这是首个流式全视频 AD 生成方法,有助于推动可访问性扩展。
正文节选
StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos Abstract Visual content is the dominant medium of communication, yet without audio descriptions (ADs), it remains inaccessible to blind and low-vision people. ADs narrate context-relevant visual events during natural audio pauses. Manually creating ADs is expensive, limiting coverage to a small fraction of available content. Most existing automatic AD generation methods frame the task as video clip capt