返回全部动态

视觉语言模型能否评估机器人自我中心图像的邻近风险?

原标题:Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?

arXiv cs.CV一手来源研究质量 81

AI 摘要

本研究评估了三种开源视觉语言模型(InternVL、Qwen-VL和SmolVLM)在机器人自我中心图像中评估邻近风险的能力,使用基于JRDB构建的1243张图像数据集,比较三种提示策略和两轮QLoRA微调。结果显示,未经微调的模型性能接近基线,微调仅带来适度提升,但Qwen-VL在高级提示下对高风险案例的召回率显著更高。分析还发现,正确的危险分类并不对应更好的空间定位,表明当前VLM在细粒度邻近推理和空间定位方面仍有限。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images? Abstract Assessing proxemic danger from a robot’s egocentric perspective is critical for safe embodied navigation in human environments and requires both visual and contextual reasoning. We evaluate three opensource vision-language models (VLMs) (InternVL, Qwen-VL, and SmolVLM) on the classification of egocentric robot images into four danger levels, comparing three prompting strategies and two rounds of QLoRA fine-tun


发布时间:2026-08-14 12:00
抓取时间:2026-08-14 13:05
来源机构:arXiv
阅读原文arxiv.org