返回全部动态
FATE:帧级音视频时间嵌入方法
原标题:FATE: Frame-Level Audio-Visual Temporal Embedding
AI 摘要
FATE是一种帧级音视频时间嵌入方法,旨在同时捕捉语义和时间对齐。与现有嵌入模型和同步模型不同,FATE保留帧级序列并在物理时间线上对齐,通过联合目标训练。在三个任务上,FATE在时间和语义检索上大幅超越最强基线,在零样本事件定位上匹配全监督方法,并作为生成评估指标与人类判断相关性最佳。源代码已公开。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
FATE: Frame-Level Audio-Visual Temporal Embedding Abstract When a dog opens its mouth and barks, humans naturally recognize what the sound is and when it occurs. Building audio-visual models with this same ability requires representations that capture both semantic and temporal alignment. Current approaches fall short on one side or the other: embedding models match semantic but lose temporal information; synchronization models capture temporal offsets but lack semantic understanding. To bridge
发布时间:—
抓取时间:2026-08-10 19:04
来源机构:Hugging Face