返回全部动态

打破视觉-动作捷径:面向可泛化机器人基础模型的潜在接口训练

原标题:Paper page - Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

该论文提出 Latent Interface Training(LIT),一种与框架无关的两阶段训练策略,用于提升机器人基础模型在视觉分布偏移下的动作泛化能力。第一阶段在不使用图像的情况下,基于语言、机器人状态和末端执行器 SE(3) 位姿训练动作先验;第二阶段引入受位姿监督的潜在接口,作为预训练动作专家唯一的视觉条件通路。在 Pi0.5、MolmoAct2、FAST-WAM 和 ImageWAM 四种架构上,LIT 使 LIBERO-Plus 成功率提升 3.87-10.70 个百分点,真实世界任务在未见相机配置、光照变化和干扰物下成功率提升 13.30-16.70 个百分点。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models Abstract LIT improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information. Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pre


发布时间:—
抓取时间:2026-09-14 10:58
来源机构:Hugging Face
阅读原文huggingface.co