返回全部动态
VidParse:在线解析自我中心程序,无需训练
原标题:VidParse: Online Parsing of Egocentric Procedures Like a Pro
AI 摘要
VidParse 是一个无需训练的在线自我中心程序解析框架,利用冻结的 DINOv2 和手-物检测器构建操作锚定特征,通过时间相似性矩阵和棋盘核检测动作边界,并使用图约束波束搜索强制执行程序性任务图。在 GTEA 和 EgoPER 基准上,该方法在复杂多步解析准确率上比强在线基线提升高达 10 倍,且无需梯度更新。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
VidParse: Online Parsing of Egocentric Procedures Like a Pro Abstract Translating continuous, noisy egocentric video streams into discrete, temporally ordered action steps is fraught with visual challenges. Heavy ego-motion, transient occlusions, and the high intra-class variability of unscripted human-object interactions cause standard frame-level online temporal models to struggle, often resulting in severe over-segmentation and structural collapse. To bridge the gap between unstable low-level
发布时间:2026-08-31 12:00
抓取时间:2026-08-31 13:01
来源机构:arXiv