返回全部动态

VibeWorlding:多模态智能体端到端构建 3D 开放世界的统一框架

原标题:Paper page - VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

VibeWorlding 是一个用于基准测试和训练多模态智能体的统一框架,旨在从用户查询端到端构建交互式 3D 开放世界。该框架包含 VWE-BENCH 基准(含 2,616 个 3D 资产、323 个种子世界和 6,828 个查询)和 VibeWorlding-Gym 强化学习训练环境。实验表明,当前前沿多模态大模型(如 GPT-5.5 和 Qwen3.8-Max)的成功率低于 60%,而通过强化学习训练的开源模型 VibeWorlder-30B-A3B 在 Pass@1 上超越了闭源模型。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? Abstract A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers. Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systemati


发布时间:—
抓取时间:2026-08-18 10:34
来源机构:Hugging Face
阅读原文huggingface.co