返回全部动态

VidParse:在线解析自我中心程序,无需训练

原标题:VidParse: Online Parsing of Egocentric Procedures Like a Pro

arXiv cs.CV一手来源研究质量 82

AI 摘要

VidParse 是一个无需训练的在线自我中心程序解析框架,利用冻结的 DINOv2 和手-物检测器构建操作锚定特征,通过时间相似性矩阵和棋盘核检测动作边界,并使用图约束波束搜索强制执行程序性任务图。在 GTEA 和 EgoPER 基准上,该方法在复杂多步解析准确率上比强在线基线提升高达 10 倍,且无需梯度更新。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

VidParse: Online Parsing of Egocentric Procedures Like a Pro Abstract Translating continuous, noisy egocentric video streams into discrete, temporally ordered action steps is fraught with visual challenges. Heavy ego-motion, transient occlusions, and the high intra-class variability of unscripted human-object interactions cause standard frame-level online temporal models to struggle, often resulting in severe over-segmentation and structural collapse. To bridge the gap between unstable low-level


发布时间:2026-08-31 12:00
抓取时间:2026-08-31 13:01
来源机构:arXiv
阅读原文arxiv.org