返回全部动态

当 Agent 指标测量不同事物:Praxa AI 管线的证据审计

原标题:When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline

arXiv cs.SE一手来源研究质量 78

AI 摘要

该论文对 Praxa AI 的 agent 评估管线进行回溯性测量审计,发现多项指标在数值正确的情况下测量了与标签不符的构念:139 例离线路由报告有 112 通过、27 失败,但门控失败为零,因为已知缺口被显式豁免;8,843 条工具尝试记录中 448 条时长缺失,且 121 条时长等于 32 位有符号整数最大值并带有被放弃客户端标签;单轨迹压缩试点报告的后续输入减少 94.39%,但包含触发调用后仅减少 46.54%。作者复现了描述性计算、用独立加权有理算术验证 91 项计时统计,并执行了 13 个评分函数测试和 12 个分析验证器测试,强调历史 provider 运行与完整当前管线未被独立复现,不主张通用能力优越性或总体统计显著性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline Abstract Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels. We present a retrospective measurement audit of selected Praxa AI implementation files, historical evaluation artifacts, and operational records. A 139-case offline routing report contains 112 passes and 27 failures despite zero gating failures, because known gaps are expl


发布时间:2026-09-14 12:00
抓取时间:2026-09-14 13:54
来源机构:arXiv
阅读原文arxiv.org