VibeLifeBench:评估生活代理在动态世界中的主动性与持久性
原标题:VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
AI 摘要
Hugging Face 发布新基准 VibeLifeBench,用于评估长期主动型生活代理,包含 200 个跨 10 个日常领域的多周任务,模拟 22 个模拟服务和 288 个工具。评估发现七个前沿模型表现均不佳,最佳模型 Claude Opus 5 平均得分仅 32.5,且所有模型在时间线推进中性能下降 10-15 分。该基准强调代理需主动发现静默变化并遵守未明示约束,任务、环境和评估框架将开源。
正文节选
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? Abstract A new benchmark called VibeLifeBench evaluates long-horizon proactive agents across simulated multi-week everyday tasks, revealing that current frontier models perform poorly. Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs f