返回全部动态

Puffin-World:原生3D世界状态驱动的统一多模态模型

原标题:Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Hugging Face Blog一手来源研究质量 90

AI 摘要

Hugging Face 博客介绍了 Puffin-World,一个统一的多模态世界模型,通过原生表示物理(重力场和纬度)、几何(深度)和外观(RGB)状态,实现相机到世界的理解、可控文本到图像生成、3D 世界生成和重建。该模型采用 Omni-Camera 表示,结合绝对和相对相机信息,并利用 Puffin-16M 数据集进行扩展,旨在推动多模态模型从 2D 语义向物理基础的 3D 世界发展。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

📖 Project Page | 💻 GitHub | 🤗 Models | 🗂️ Dataset | 🧪 Supplementary A world model should do more than generate plausible pixels. It should know where the camera is in the physical world, understand the geometry beneath an observation, and generate new viewpoints while remaining consistent with gravity and scene structure. We introduce Puffin-World, a unified multimodal world model that perceives, simulates, generates, and reconstructs the 3D world within one framework. Instead of treating a worl


发布时间:—
抓取时间:2026-09-03 14:36
来源机构:Hugging Face
阅读原文huggingface.co