返回全部动态

HappyWorld-Bench:评估世界模型交互可靠性的综合基准

原标题:Paper page - HappyWorld-Bench

Hugging Face Daily Papers一手来源研究质量 81

AI 摘要

研究者提出 HappyWorld-Bench,一个用于评估世界模型在智能体交互、探索与修改下是否保持可靠的综合基准。该基准基于六项世界能力(W1-W6)的分层框架,覆盖视频、空间与具身三条评测赛道,包含 1,138 条视频提示、300 个空间场景和 254 个具身测试用例,并搭建 HappyWorld-Arena 进行人类 A/B 对比与 Elo 评分。评测 14 个视频世界模型、9 个空间系统和 8 个具身候选模型后发现,三条赛道均存在可靠性缺口:视频模型在长时推演与重访时一致性下降,空间模型最高仅 70.14% 放置准确率和 73.33% 编辑执行率,具身模型难以在多步动作中保持状态。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

HappyWorld-Bench Abstract Evaluating world models requires assessing both the quality of the worlds they generate and their consistency and responsiveness under exploration, interaction, and modification. We introduce HappyWorld-Bench, a comprehensive benchmark that evaluates whether generated worlds remain reliable as agents interact with them. Our design is built on a hierarchical capability framework of six world capabilities (W1-W6), from generative construction to unified world modeling, in


发布时间:—
抓取时间:2026-09-24 16:00
来源机构:Hugging Face
阅读原文huggingface.co