Harness还是模型?在防污染私有套件中隔离智能体编码的Harness效应
原标题:Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite
AI 摘要
该研究通过配对同模型对比,在256个私有防污染任务上测试了厂商原生harness与中立harness(deepagents)对智能体编码能力的影响。结果显示,Opus 4.8和GPT-5.5上两种harness的解决率差异均不显著(分别约1.25个百分点),但Opus在仓库任务上原生harness落后9.0个百分点、在竞赛任务上领先23.7个百分点,呈现任务类型交互。成本方面,中立harness每解决一个任务的成本更高(Opus为1.3-1.6倍,GPT-5.5为1.2倍),但Anthropic账户存在未记录用量导致成本排序未定。研究还修正了此前版本中因遥测用量语义缺陷导致的成本结论。
正文节选
Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private SuiteNote: Revised 8 September 2026. A post-study review found a usage-semantics defect in the study’s cost telemetry. Relative to the August 2026 manuscript, the cost results, the probe interpretation and the registration statement are corrected, and the capability analysis is re-run with the protocol’s tests. The revision memo and reanalysis code are in the replication package. Abstract. An