返回全部动态

轨迹裁判:仅看结果的 LLM 裁判在智能体轨迹上的盲区

原标题:trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

arXiv cs.CL一手来源研究质量 87

AI 摘要

本研究构建了一个确定性工具使用环境,通过脚本化 oracle 策略和故障注入器,对五种 LLM 裁判(包括仅看结果的裁判、步骤评分裁判等)在 400 条轨迹上的表现进行了评估。结果显示,仅看结果的裁判对静默故障的召回率仅为 45%,且会误报 33% 的正确轨迹;而步骤评分裁判在零误报的情况下达到了 77% 的静默召回率。研究还发现,自一致性集成成本增加三倍但无改进,且所有裁判都无法识别虚构承诺。作者主张裁判评估应按结果存活率分层,并发布了完整可复现的测试环境。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories Abstract Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well. The metric is structurally blind to an agent that reaches the right answer the wrong way. We measure that blind spot where ground truth is known by construction: a deterministic tool-using support-desk environment, a scripted oracle policy that always solves it, and a


发布时间:2026-09-02 12:00
抓取时间:2026-09-02 12:10
来源机构:arXiv
阅读原文arxiv.org