返回全部动态

HarnessEval-W:将世界模型评估代理化

原标题:Paper page - HarnessEval-W: Agentifying the Evaluation of Visual Worlds

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

HarnessEval-W 是一个用于世界模型评估的代理化流水线,通过分层子代理将评估分解为可验证的推理链,并以透明证据证明评分。该流水线应用于18个世界模型的330个评估案例,其判断与人类偏好高度一致,同时提供可验证的细粒度诊断。作者已开源完整流水线作为实时基准,邀请社区贡献新的技能和评估案例。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

HarnessEval-W: Agentifying the Evaluation of Visual Worlds Abstract HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence. A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state ev


发布时间:—
抓取时间:2026-08-18 16:56
来源机构:Hugging Face
阅读原文huggingface.co