评估企业分析智能体:端到端、轨迹支撑的方法论
原标题:Evaluating Enterprise Analytics Agents: An End-to-End, Trace-Backed Methodology
AI 摘要
该论文提出了一套面向企业分析智能体的端到端、基于运行轨迹的评估方法,从语义理解、执行质量和可靠性三个维度对智能体行为进行评分,而非仅看最终答案或SQL正确性。作者在一个大型在线市场内部署的受控分析智能体上进行了案例研究,使用50个问题、两种匿名模型配置和三次随机重复,共获得300条轨迹。结果显示,能力更强的配置将早期拒绝率从73%降至0%,真实数据回答率从21%升至73%,但在16%的运行中耗尽工具轮次预算,在77%的轨迹中超出模式探索预算,并在50个问题中的41个上改变了表解释。
正文节选
Evaluating Enterprise Analytics Agents: An End-to-End, Trace-Backed Methodology Abstract. Enterprise analytics agents are not only text-to-SQL systems. They interpret business intent and choose metric definitions. They select data sources, execute tools, inspect results, and produce natural-language answers. Those answers may influence operational, financial, or executive decisions. Grading final answers hides where these agents fail. A plausible answer can use the wrong source of truth. It can