当 Agent 指标测量不同事物:Praxa AI 管线的证据审计
原标题:When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline
AI 摘要
该论文对 Praxa AI 的 agent 评估管线进行回溯性测量审计,发现多项指标在数值正确的情况下测量了与标签不符的构念:139 例离线路由报告有 112 通过、27 失败,但门控失败为零,因为已知缺口被显式豁免;8,843 条工具尝试记录中 448 条时长缺失,且 121 条时长等于 32 位有符号整数最大值并带有被放弃客户端标签;单轨迹压缩试点报告的后续输入减少 94.39%,但包含触发调用后仅减少 46.54%。作者复现了描述性计算、用独立加权有理算术验证 91 项计时统计,并执行了 13 个评分函数测试和 12 个分析验证器测试,强调历史 provider 运行与完整当前管线未被独立复现,不主张通用能力优越性或总体统计显著性。
正文节选
When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline Abstract Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels. We present a retrospective measurement audit of selected Praxa AI implementation files, historical evaluation artifacts, and operational records. A 139-case offline routing report contains 112 passes and 27 failures despite zero gating failures, because known gaps are expl