返回全部动态

全失败任务真的难吗?区分真实难度与假难度

原标题:What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

arXiv cs.LG一手来源研究质量 83

AI 摘要

该论文基于冻结的 Terminal-Bench 3 / Frontier-Bench 0.1 生产记录(1,081 个 PR、639 个计分任务、28,801 次试验、约 10.6 万美元 agent 支出),研究全失败任务是否真正代表能力缺口。作者对 125 个无诚实通过的任务施加有序有效性筛查,仅 78 个被认证为「未解决候选」,其余包括 14 个 oracle 损坏、8 个基础设施主导、4 个仅能通过验证器绕过通过、21 个可解性未获证据认证。结论指出低通过率不等于真实难度,前沿基准应报告全失败任务背后的证据再将其用作能力声明。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus Abstract Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be bypassed. In this paper, we study this issue using a froze


发布时间:2026-09-24 12:00
抓取时间:2026-09-24 12:15
来源机构:arXiv
阅读原文arxiv.org