W2-VLA:任务条件化未来手腕建模实现精细机器人操作
原标题:World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
AI 摘要
Hugging Face 发布了 W2-VLA 模型,该模型通过任务条件化的未来手腕建模来提升精细机器人操作能力。W2-VLA 利用潜在建模令牌作为视觉语言模型与手腕预测器之间的接口,并引入 W2-CoT 合成管道生成结构化注释以辅助训练。实验表明,在 LIBERO、RoboTwin 2.0 和真实世界任务中,该模型在单臂和双臂场景下均取得了改进,动作生成速率超过 80 Hz。
正文节选
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation Abstract Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained