WorldCrafter:带隐式3D感知记忆的一致视频世界模型
原标题:Paper page - WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
AI 摘要
研究团队提出 WorldCrafter,一种视频世界模型,通过可相机查询的隐式 3D 感知记忆来提升长时程与跨视角一致性。其核心思路是让请求视角决定多视角证据如何压缩进视频生成器有限的 token 预算,记忆编码器与姿态条件读出模块与生成器联合训练,无需显式深度对应。结合近期时序上下文与少步蒸馏,模型可从单张图像或文本提示进行流式场景探索,在静态与动态场景中显著提升长时程一致性与相机控制精度。
正文节选
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Abstract Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budg