返回全部动态

VLM 可供性预测瓶颈在于部件定位而非动作知识

原标题:Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction

arXiv cs.CV一手来源研究质量 78

AI 摘要

该研究将视觉语言模型(VLM)的 affordance 预测拆分为「定位操作部件」与「知道该部件所需动作」两步,在 19 个铰接物体上测试了来自三家开发商的八个模型。开放提示下 push 动作在 64 次评估中仅正确出现一次,八个模型中有七个从未输出 push,表面看像是动作知识缺失。但检查输出发现模型实际在描述与评分不同的部件,在提示中明确命名目标部件后,所有模型动作准确率从 0.158–0.474 提升至 0.684–0.947,push 召回率从 0–1/8 升至 7–8/8。结论是部件定位而非动作知识才是主要瓶颈,且该模式跨三个模型家族且不随模型能力增强而减弱。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Part Grounding, Not Action Knowledge: Locating the Bottleneck in VLM Affordance Prediction Abstract Benchmarks agree that vision-language models reason poorly about low-level manipulation, but an aggregate accuracy does not say which step fails. We separate two steps that affordance questions usually conflate: identifying which part of an object is the one to act on, and knowing what action that part requires. On 19 articulated objects we asked eight models, spanning three developers, what motio


发布时间:2026-09-15 12:00
抓取时间:2026-09-15 12:22
来源机构:arXiv
阅读原文arxiv.org