返回全部动态

超越最终分数:长时程AI研发智能体的系统评估

原标题:Paper page - Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

该论文提出了一种系统评估方法,对7个前沿模型在36个长时程任务上的表现进行了评估,超越了最终分数,通过基于规则的指标分析运行内行为和经验复用。研究发现,当前智能体更像工程优化器而非完全自主的研究者,性能不稳定,创新性有限,经验复用可能有益也可能误导后续决策。研究还指出,相似的最终分数可能掩盖不同的过程瓶颈,且框架设计显著影响性能稳定性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Abstract Frontier autonomous agents excel at engineering optimization but show unstable performance, limited novelty, and variable experience reuse across long-horizon tasks. Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyo


发布时间:—
抓取时间:2026-08-17 10:47
来源机构:Hugging Face
阅读原文huggingface.co