视觉语言模型能否评估机器人自我中心图像的邻近风险?
原标题:Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?
AI 摘要
本研究评估了三种开源视觉语言模型(InternVL、Qwen-VL和SmolVLM)在机器人自我中心图像中评估邻近风险的能力,使用基于JRDB构建的1243张图像数据集,比较三种提示策略和两轮QLoRA微调。结果显示,未经微调的模型性能接近基线,微调仅带来适度提升,但Qwen-VL在高级提示下对高风险案例的召回率显著更高。分析还发现,正确的危险分类并不对应更好的空间定位,表明当前VLM在细粒度邻近推理和空间定位方面仍有限。
正文节选
Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Abstract Assessing proxemic danger from a robot’s egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (InternVL, Qwen-VL, and SmolVLM) on the classification of egocentric robot images into four danger levels, comparing three prompting strategies and two rounds of QLoRA fine-tun