RLHEV:利用游戏开发作为可验证轨迹数据引擎扩展世界模型
原标题:Paper page - Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models
AI 摘要
该论文提出了一种名为RLHEV(Reinforcement Learning with Human-Engine Verification)的后训练范式,利用游戏引擎作为可验证的轨迹数据引擎,为空间世界模型的强化学习后训练提供密集的奖励信号和长时程轨迹数据。作者认为,仅通过增加数据和计算量来扩展世界模型效率低下,而游戏开发提供了可执行的环境,能够提供碰撞、物理、可导航性等验证信号,并结合开发者的隐式反馈,从而支持更有效的RL后训练。
正文节选
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Abstract Game engines provide executable verification and long-horizon trajectories for reinforcement learning post-training of spatial world models, motivating a human-engine verification paradigm. A common strategy for scaling world models is to train on more crawled video with more compute. We argue that this strategy is inefficient: scaling world models also requires a recursive data engine that offers g