返回全部动态

Harness还是模型?在防污染私有套件中隔离智能体编码的Harness效应

原标题:Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

arXiv cs.AI一手来源研究质量 85

AI 摘要

该研究通过配对同模型对比,在256个私有防污染任务上测试了厂商原生harness与中立harness(deepagents)对智能体编码能力的影响。结果显示,Opus 4.8和GPT-5.5上两种harness的解决率差异均不显著(分别约1.25个百分点),但Opus在仓库任务上原生harness落后9.0个百分点、在竞赛任务上领先23.7个百分点,呈现任务类型交互。成本方面,中立harness每解决一个任务的成本更高(Opus为1.3-1.6倍,GPT-5.5为1.2倍),但Anthropic账户存在未记录用量导致成本排序未定。研究还修正了此前版本中因遥测用量语义缺陷导致的成本结论。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private SuiteNote: Revised 8 September 2026. A post-study review found a usage-semantics defect in the study’s cost telemetry. Relative to the August 2026 manuscript, the cost results, the probe interpretation and the registration statement are corrected, and the capability analysis is re-run with the protocol’s tests. The revision memo and reanalysis code are in the replication package. Abstract. An


发布时间:2026-09-14 12:00
抓取时间:2026-09-14 12:14
来源机构:arXiv
阅读原文arxiv.org