返回全部动态

智能体基准测试的双重测量混淆:去脚手架、真值评分与超越均值的可靠性

原标题:The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean

arXiv cs.SE一手来源研究质量 90

AI 摘要

该论文指出LLM智能体基准测试存在双重测量混淆:执行关键决策由固定脚手架而非模型完成,评分器又基于输出形状而非任务正确性打分,两者相互掩盖。作者提出测量理论框架与BenchAudit审计修复协议,将执行决策转移给模型并用种子化真值评分,同时报告最坏情况和尾部风险等可靠性指标。在ComtradeBench上的实验将原本几乎无区分度的排行榜转变为能区分平均性能与跨种子鲁棒性的可靠性谱,外部审计还发现官方评分器受脚手架影响可改变模型排名。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean Abstract Agent benchmarks are increasingly used to compare large language models (LLMs) and guide deployment decisions, yet benchmark scores are meaningful only if they measure model capability rather than properties of the evaluation pipeline. We identify a double measurement confound: execution-critical decisions are performed by a fixed scaffold instead of the model, whil


发布时间:2026-09-10 12:00
抓取时间:2026-09-10 13:04
来源机构:arXiv
阅读原文arxiv.org