发表机构
Huawei Ireland Research Center; Huawei Shenzhen R&D Center(华为爱尔兰研究中心; 华为深圳研发中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生产环境智能体轨迹评估成本高的问题,提出轻量级架构LiteTrajEval,通过离线规则配置和在线预处理,结合单一评分标准LLM评判器,在显著降低成本和评估时间的同时,大幅提升失败定位与人工标注的一致性,并已部署于企业平台。
AI 中文摘要
轨迹评估对于提升基于大语言模型(LLM)的智能体的可靠性至关重要,但在生产环境中反复运行会带来高昂的成本。现代智能体生成包含工具调用、观察、重试和外部输出的长轨迹,而并非所有原始标记(token)对诊断同等有用。我们提出LiteTrajEval,一种用于预算受限轨迹评估的轻量级架构。LiteTrajEval离线推导出紧凑的领域特定规则配置文件,然后在线预处理每条轨迹,标记启发式失败信号,在固定的全局预算下将其序列化,并调用单一的基于评分标准的LLM评判器生成结构化诊断报告。在公开的Magentic-One风格和τ-bench风格轨迹数据集上评估,与AgentRx相比,LiteTrajEval在Magentic-One上将失败定位与人工标注的一致性提高了约20至35个百分点,在τ-retail上提高了最多23个百分点,同时成本降低约6倍,评估时间减少超过8倍。该解决方案也已在我们的企业智能体平台中部署。
英文摘要
Trajectory evaluation is essential for improving the reliability of LLM-based agents, but production use makes it expensive to run repeatedly. Modern agents generate long traces containing tool calls, observations, retries, and external outputs, while not all raw tokens are equally useful for diagnosis. We present \textit{LiteTrajEval}, a lightweight architecture for budget-bounded trajectory evaluation. LiteTrajEval derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports. Evaluated on public Magentic-One-style and $τ$-bench-style trajectory datasets, LiteTrajEval improves failure-localization alignment with human annotations by roughly 20--35 percentage points on Magentic-One and up to 23 percentage points on $τ$-retail compared with AgentRx, while reducing cost by about 6$\times$ and evaluation time by more than 8$\times$. This solution has also been deployed in our enterprise agentic platform.