打破视觉-动作捷径:面向可泛化机器人基础模型的潜在接口训练
原标题:Paper page - Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
AI 摘要
该论文提出 Latent Interface Training(LIT),一种与框架无关的两阶段训练策略,用于提升机器人基础模型在视觉分布偏移下的动作泛化能力。第一阶段在不使用图像的情况下,基于语言、机器人状态和末端执行器 SE(3) 位姿训练动作先验;第二阶段引入受位姿监督的潜在接口,作为预训练动作专家唯一的视觉条件通路。在 Pi0.5、MolmoAct2、FAST-WAM 和 ImageWAM 四种架构上,LIT 使 LIBERO-Plus 成功率提升 3.87-10.70 个百分点,真实世界任务在未见相机配置、光照变化和干扰物下成功率提升 13.30-16.70 个百分点。
正文节选
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models Abstract LIT improves robot action generalization by first training pose-conditioned action priors without images, then constraining visual inputs through a pose-supervised latent interface that preserves spatial goal information. Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pre