Google 研究:多智能体框架自动生成连贯长视频
原标题:Automating coherent long-form video generation
AI 摘要
Google Research 提出一个统一的多智能体框架,作为 Gemini 和 Veo 之上的编排层,自动生成时间一致的长篇视频叙事,缓解现有线性流水线中的身份漂移和级联失败问题。该框架包含 Co-Director、CANVAS、A²RD 和 VQQA 等组件,将长篇视频生成建模为全局优化与世界状态跟踪问题,并原生继承 SynthID 水印等安全机制。评估显示其在多镜头叙事一致性和角色持久性上有显著提升,可生成数分钟长的视频。
正文节选
September 24, 2026 Yale Song and Yiwen Song, Research Scientists, Google We introduce a unified multi-agent framework that autonomously generates temporally consistent, long-form video narratives, overcoming the identity drift and cascading failures of current linear AI pipelines. Recent advancements in video diffusion demonstrate remarkable high-fidelity generation with models that can render realistic scenes in seconds. However, while diffusion models generate high-fidelity video clips, transf