超越最终分数:长时程AI研发智能体的系统评估
原标题:Paper page - Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
AI 摘要
该论文提出了一种系统评估方法,对7个前沿模型在36个长时程任务上的表现进行了评估,超越了最终分数,通过基于规则的指标分析运行内行为和经验复用。研究发现,当前智能体更像工程优化器而非完全自主的研究者,性能不稳定,创新性有限,经验复用可能有益也可能误导后续决策。研究还指出,相似的最终分数可能掩盖不同的过程瓶颈,且框架设计显著影响性能稳定性。
正文节选
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Abstract Frontier autonomous agents excel at engineering optimization but show unstable performance, limited novelty, and variable experience reuse across long-horizon tasks. Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyo