返回全部动态

EvoHarnessBench:评估智能体在演化 harness 下的适应能力

原标题:EVOHARNESSBENCH: Can Your Agents Keep Pace with an Evolving Harness?

arXiv cs.MA一手来源研究质量 84

AI 摘要

EvoHarnessBench 是一个新基准,用于评估 LLM 智能体在工具、技能和智能体三个维度上不断演化的外部 harness 下的表现。它包含 17 个多阶段流、802 个任务、520 个工具、42 个技能和 62 个智能体,并区分部署评估和自进化适应评估。研究发现,harness 扩展会导致智能体在已解决任务上性能下降,即 harness 诱导的遗忘,且保留与适应之间存在权衡。该基准揭示了 harness 演化是构建持久智能体的独特挑战。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness? Abstract Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EvoHarnessBench, a benchmark for evaluating agents under controlled harness evolution across three axes (tools, skills, and agents). Unlike existing continual-learning benchmarks f


发布时间:2026-09-07 12:00
抓取时间:2026-09-07 12:52
来源机构:arXiv
阅读原文arxiv.org