返回全部动态

NVIDIA 提出世界动作模型,Cosmos 3 助力机器人操作泛化

原标题:Beyond VLAs: How World Action Models Reshape Robot Manipulation

NVIDIA Technical Blog一手来源研究质量 84

AI 摘要

NVIDIA 技术博客探讨了从视觉-语言-动作模型(VLA)向世界动作模型(WAM)的转变,指出 WAM 基于视频世界模型构建,能更好地泛化物理交互。NVIDIA 发布了 Cosmos 3 世界基础模型,并基于其训练了 DROID 机器人策略,实验显示 omni 检查点将成功率从 28.1% 提升至 36.8%。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

A central challenge in robotics is building policies that generalize beyond the demonstrations they’re trained on. A policy that succeeds in a training scene often fails when object shapes, positions, or lighting change. Generalizing to these new conditions requires the policy to understand the tasks underlying physics, not just mimic the demonstrations. This ability comes from the backbone it’s built on. The standard way to build a language-conditioned robot policy is to add an action module to


发布时间:2026-08-05 00:00
抓取时间:2026-08-05 00:23
来源机构:NVIDIA
阅读原文developer.nvidia.com