返回全部动态
HarnessEval-W:将世界模型评估代理化
原标题:Paper page - HarnessEval-W: Agentifying the Evaluation of Visual Worlds
AI 摘要
HarnessEval-W 是一个用于世界模型评估的代理化流水线,通过分层子代理将评估分解为可验证的推理链,并以透明证据证明评分。该流水线应用于18个世界模型的330个评估案例,其判断与人类偏好高度一致,同时提供可验证的细粒度诊断。作者已开源完整流水线作为实时基准,邀请社区贡献新的技能和评估案例。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Abstract HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence. A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state ev
发布时间:—
抓取时间:2026-08-18 16:56
来源机构:Hugging Face