Puffin-World:原生3D世界状态驱动的统一多模态模型
原标题:Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
AI 摘要
Hugging Face 博客介绍了 Puffin-World,一个统一的多模态世界模型,通过原生表示物理(重力场和纬度)、几何(深度)和外观(RGB)状态,实现相机到世界的理解、可控文本到图像生成、3D 世界生成和重建。该模型采用 Omni-Camera 表示,结合绝对和相对相机信息,并利用 Puffin-16M 数据集进行扩展,旨在推动多模态模型从 2D 语义向物理基础的 3D 世界发展。
正文节选
📖 Project Page | 💻 GitHub | 🤗 Models | 🗂️ Dataset | 🧪 Supplementary A world model should do more than generate plausible pixels. It should know where the camera is in the physical world, understand the geometry beneath an observation, and generate new viewpoints while remaining consistent with gravity and scene structure. We introduce Puffin-World, a unified multimodal world model that perceives, simulates, generates, and reconstructs the 3D world within one framework. Instead of treating a worl