RynnValue:利用时间距离扩展机器人价值基础模型
原标题:RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
AI 摘要
阿里巴巴达摩院开源了机器人操作价值基础模型RynnValue,利用时间距离作为监督信号,在超过7000小时、约300万条指令条件片段的数据上训练,无需偏好或进度标注。该模型在RBM-EVAL-OOD基准上平均Kendall's tau_a达到0.675,超过完全偏好监督的现有最优方法(0.655),并显著优于仅使用进度的方法(0.292)。通过基于势能的奖励塑形,RynnValue将真实世界策略成功率从52.5%提升至72.5%(在线)和从63.8%提升至82.5%(离线),展示了时间距离作为可扩展监督目标和实用奖励接口的有效性。
正文节选
RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance Abstract General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-sourc