返回全部动态

ZimaBlue:通过可扩展视频预训练进化通用世界动作模型

原标题:ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

arXiv cs.CV一手来源研究质量 88

AI 摘要

ZimaBlue 是一个从大规模视频中学习通用世界动作模型(WAM)的可扩展框架,采用三阶段训练课程:因果具身视频预训练、视频-动作中间训练和针对目标机器人的微调。其异步慢-快双系统架构支持在 RTX 4090 上实现 30 Hz 的实时动作预测。在真实机器人零样本评估中,将训练数据从仅目标机器人数据扩展到超过 12 万小时的具身视频,成功率从 36.1% 提升至 77.8%,并在多个基准上表现优异,尤其在未见任务上提升显著。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

fadings ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training Abstract Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Egocentric videos offer a far more scalable source of embodied experience, capturing object interactions, contact dynamics, tool use, and long-horizon behaviors across diverse e


发布时间:2026-09-02 12:00
抓取时间:2026-09-02 12:35
来源机构:arXiv
阅读原文arxiv.org