返回全部动态

验证工具覆盖面决定AI编码智能体价值:一项受控研究

原标题:The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents

arXiv cs.SE一手来源研究质量 84

AI 摘要

该研究通过构建一个工具列表为唯一变量的最小化编码智能体,在六个模型和八种工具配置下实现了1116个Web应用,以探究验证工具覆盖面(如linter、启动探针、shell、截图)对AI编码智能体输出软件质量的影响。研究发现,验证工具的最廉价收益是确保应用能启动,仅一个启动探针即可消除大部分启动失败,且成本仅为完整shell的35%;完整shell将无工具成本提高2.35倍,而截图仅在可见错误(如元素布局)上有微小提升,对仅可测量的失败(如滚动性能)无帮助。结论是验证工具仅在覆盖实际失败模式时才能改善输出。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

[Page 1] The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents Achint Mehta Corresponding author: Achint Mehta (e-mail: achintmehta@gmail.com) This work received no external funding. ABSTRACT Modern artificial-intelligence coding agents can be equipped with tools for checking their own work e.g. a linter, a boot probe, a shell, a screenshot tool. We call this set the agent's verification surface. This s


发布时间:2026-09-01 12:00
抓取时间:2026-09-01 14:12
来源机构:arXiv
阅读原文arxiv.org