ATLAS:面向工业工具使用智能体的双视角诊断评估框架
ATLAS: Dual-Horizon Diagnostic Evaluation for Industrial Tool-Use Agents
- University of Science and Technology of China(中国科学技术大学)
- Meituan(美团)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对工业工具使用LLM智能体的评估缺陷,提出双视角诊断评估框架ATLAS,经美团小团生产流量验证,可提升服务质量并优化策略。
AI中文摘要:
大型语言模型(LLM)智能体正越来越多地部署在需在动态业务条件下进行迭代工具使用的面向用户服务中。可靠的评估是持续改进的关键:它必须揭示能力缺陷、明确优先级并评估干预措施。然而,工业智能体服务既通过当前请求的迭代轨迹展开,也通过持续的用户交互展开。因此,仅基于最终结果的评估会模糊缺陷产生的位置,以及后续服务是否仍与早期交互的上下文保持一致。我们提出了ATLAS,这是一个面向工业工具使用智能体的双视角诊断评估框架。在请求视角下,基于轨迹的诊断信号将缺陷与执行位置和能力问题关联起来;在交互视角下,基于用户的信号评估服务在持续交互中是否仍保持响应性。这些视角共同为分析执行缺陷和持续服务行为提供了结构化诊断证据。ATLAS将这些视角实例化为具有明确证据范围和决策边界的可执行信号。LLM评估接口针对来自真实业务日志的高置信度参考进行校准;必要时,其决策行为会被提炼为高效诊断模型,以实现更低延迟、更低成本的评估。生成的反馈支持策略优化。我们在美团小团的生产流量上对ATLAS进行了评估。离线实验评估了诊断信号的保真度和基于重放的策略改进,而在线A/B实验显示用户参与度、下游业务结果和抽样人工审计质量均获得了同步提升。
英文摘要:
Large language model (LLM) agents are increasingly deployed in user-facing services that require iterative tool use under dynamic business conditions. Reliable evaluation is essential for sustained improvement: it must reveal capability deficiencies, inform priorities, and assess interventions. Yet industrial agent service unfolds both through the iterative trajectory of a current request and through continued user interaction. Final-outcome assessment can therefore obscure where deficiencies arise and whether later service remains aligned with context from earlier exchanges. We propose ATLAS, a dual-horizon diagnostic evaluation framework for industrial tool-use agents. At the request horizon, trajectory-wise diagnostic signals relate deficiencies to execution locations and capability concerns. At the interaction horizon, user-wise signals assess whether service remains responsive across continued interaction. Together, these views provide structured diagnostic evidence for analyzing execution deficiencies and sustained service behavior. ATLAS instantiates them as executable signals with explicit evidence scopes and decision boundaries. LLM judge interfaces are calibrated against high-confidence references from real business logs; when needed, their decision behavior is distilled into efficient diagnostic models for lower-latency, lower-cost evaluation. The resulting feedback supports policy optimization. We evaluate ATLAS on Meituan Xiaotuan production traffic. Offline experiments assess diagnostic-signal fidelity and replay-based policy improvement, while online A/B experiments show concurrent gains in user engagement, downstream business outcomes, and sampled human-audit quality.