返回全部动态

VibeLifeBench:评估生活代理在动态世界中的主动性与持久性

原标题:VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Hugging Face Daily Papers一手来源研究质量 86

AI 摘要

Hugging Face 发布新基准 VibeLifeBench,用于评估长期主动型生活代理,包含 200 个跨 10 个日常领域的多周任务,模拟 22 个模拟服务和 288 个工具。评估发现七个前沿模型表现均不佳,最佳模型 Claude Opus 5 平均得分仅 32.5,且所有模型在时间线推进中性能下降 10-15 分。该基准强调代理需主动发现静默变化并遵守未明示约束,任务、环境和评估框架将开源。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? Abstract A new benchmark called VibeLifeBench evaluates long-horizon proactive agents across simulated multi-week everyday tasks, revealing that current frontier models perform poorly. Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs f


发布时间:—
抓取时间:2026-08-12 11:02
来源机构:Hugging Face
阅读原文huggingface.co