评估企业分析智能体:一种端到端、基于追踪的方法论
Evaluating Enterprise Analytics Agents: An End-to-End, Trace-Backed Methodology
浏览论文内容
中文总结 AI 辅助
提出一种端到端、基于追踪的企业分析智能体评估方法论,通过三类评分和分层弃权机制,在50个问题上验证了高能力配置降低拒绝率但存在预算超支和解释不稳定问题。
中文摘要 AI 辅助
企业分析智能体不仅仅是文本到SQL系统。它们解读业务意图并选择指标定义。它们选择数据源、执行工具、检查结果并生成自然语言答案。这些答案可能影响运营、财务或高管决策。仅对最终答案评分会掩盖这些智能体失败之处。一个看似合理的答案可能使用了错误的事实来源。它可能跳过必要的分解,无依据地声称因果关系,或在多次运行中改变其表格解释。我们提出了一种针对分析智能体的端到端评估方法论。该方法论对智能体行为在三个类别中进行评分:语义理解、执行质量和可靠性。评分使用带有专家编写黄金答案的问题库、重复运行和运行时追踪。每次运行首先通过运行有效性检查,然后获得分层的、允许弃权(不执行)的评分,这些评分服务于决策框架而非发布门槛。我们在一个大型在线市场的受控内部分析智能体上实例化了该方法论。案例研究使用了50个分析问题、两个匿名模型配置和每个配置三次随机重复,共产生300条追踪记录。更高能力的配置将早期拒绝率从73%降至0%,并将基于真实数据的答案从21%提升至73%。然而,它也在16%的运行中耗尽了工具轮次预算。它在77%的追踪中超出了模式探索预算。它在50个问题中的41个上改变了其表格解释。在具有结构化黄金答案的财务问题上,源表使用和升级有所改善,但两种配置中的规范分解仍然薄弱。这些结果表明,对分析智能体的信任需要端到端、基于追踪的评估,而非仅依赖SQL正确性或最终答案质量。
英文摘要
Enterprise analytics agents are not only text-to-SQL systems. They interpret business intent and choose metric definitions. They select data sources, execute tools, inspect results, and produce natural-language answers. Those answers may influence operational, financial, or executive decisions. Grading final answers hides where these agents fail. A plausible answer can use the wrong source of truth. It can skip a required decomposition, claim causality without support, or change its table interpretation across repeated runs. We present an end-to-end evaluation methodology for analytics agents. The methodology grades agent behavior across three families: semantic understanding, execution quality, and reliability. Grading uses question banks with human-written golden answers, repeated runs, and runtime traces. Each run first passes a run-validity check, then receives tiered, abstention-aware scores that feed a decision framework rather than a release gate. We instantiate the methodology on a controlled internal analytics agent at a large online marketplace. The case study uses 50 analytics questions, two anonymous model configurations, and three randomized repetitions per configuration, yielding 300 traces. The higher-capability configuration reduced early refusal from 73% to 0% and increased real-data answers from 21% to 73%. However, it also exhausted the tool-round budget on 16% of runs. It overran the schema-exploration budget on 77% of traces. It changed its table interpretation on 41 of 50 questions. On finance questions with structured golden answers, source table use and escalation improved, but canonical decomposition remained weak in both configurations. These results show why trust in analytics agents requires end-to-end, trace-backed evaluation rather than SQL correctness or final answer quality alone.
发表机构
- Thumbtack
机构由 AI 辅助整理,请以论文原文为准。