返回全部动态

Show-Harness:仅靠 VLM 智能体即可操控机器人

原标题:Paper page - Show-Harness: Just a VLM Agent Can Play Robots

Hugging Face Daily Papers一手来源研究质量 80

AI 摘要

该论文提出 Show-Harness,一种通过离散语义动作单元将视觉语言模型(VLM)与机器人控制连接起来的具身接口,由特定于本体的解释器将语义动作确定性地映射为本地机器人动作。Show-Harness 证明了两点可行性:直接解锁闭源前沿 VLM 实现零样本机器人控制,以及仅用少量 GPU 小时微调小规模开源 VLM 实现低成本部署。作者还开发了 GUMI(GUI 操作接口),将同一语义动作空间扩展到基于 GUI 的演示收集,使人类和智能体无需专用遥操作硬件即可跨本体操控机器人。实验显示,配备 Show-Harness 的 VLM 智能体在任务、本体和环境间泛化稳健,优于代表性的智能体和 VLA 范式。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Show-Harness: Just a VLM Agent Can Play Robots Abstract Show-Harness links vision-language models to robot control via discrete semantic actions interpreted by embodiment-specific modules, enabling zero-shot and efficient fine-tuned deployment across robots and GUIs. Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play"


发布时间:—
抓取时间:2026-09-10 10:45
来源机构:Hugging Face
阅读原文huggingface.co