HappyWorld-Bench:评估世界模型交互可靠性的综合基准
原标题:Paper page - HappyWorld-Bench
AI 摘要
研究者提出 HappyWorld-Bench,一个用于评估世界模型在智能体交互、探索与修改下是否保持可靠的综合基准。该基准基于六项世界能力(W1-W6)的分层框架,覆盖视频、空间与具身三条评测赛道,包含 1,138 条视频提示、300 个空间场景和 254 个具身测试用例,并搭建 HappyWorld-Arena 进行人类 A/B 对比与 Elo 评分。评测 14 个视频世界模型、9 个空间系统和 8 个具身候选模型后发现,三条赛道均存在可靠性缺口:视频模型在长时推演与重访时一致性下降,空间模型最高仅 70.14% 放置准确率和 73.33% 编辑执行率,具身模型难以在多步动作中保持状态。
正文节选
HappyWorld-Bench Abstract Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, in