arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在生产环境中评估和改进对话智能体的框架

On Evaluating and Improving Conversational Agents in Production

Kasra Hosseini, Wen-Sen Cheng, Marco-Andrea Buchmann, Emir Mulabegovic, Weiwei Cheng

arXiv 2609.32092首次发表:更新:

发表机构

Zalando SE; Zalando(扎兰多股份公司; 扎兰多)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对生产环境中多智能体购物助手离线评估的三重障碍,提出包含评估工具、真实用户模拟和基线比较的框架,并通过生产调查验证其有效性。

AI 中文摘要

我们提出了一个在生产环境中评估和改进大规模多智能体购物助手的框架,并报告了其使用经验。对此类系统的离线评估面临三个障碍。(i)无法针对修改后的系统重放已记录的对话,因为不同的响应会改变后续的每一轮对话。(ii)未修改的系统本身在不同运行之间也会有所变化。其大语言模型组件是随机的,并且在产品搜索中,可用产品、价格以及客户的个性化信号都会发生变化。(iii)聚合质量分数结合了不同的行为,因此它们显示质量发生了变化,但无法显示是哪种行为导致了变化。我们的框架逐一解决了这些障碍。对于所报告的行为,评估工具生成有针对性的断言和固定的客户场景队列。然后,它通过基于真实用户模拟在助手的本地实例中重现该行为。模拟器不是重放日志,而是根据记录的消息和上下文生成新的客户轮次。对未修改系统的重复运行形成了存储的基线。改进协调器将断言结果转化为假设,将每个假设作为独立的修改实现,并使用基于场景级差异的配对百分位自助法区间将其与基线进行比较。当调查结束时,工具可能会提出对未来评估的修订建议,但需经人工批准,且不会改变过去的决策。我们报告了使用该框架进行的生产调查。断言配置文件显示了失败影响产品轮播的哪些位置,重复运行将真正的改进与运行间的波动区分开来。对评估本身的审计发现,一个评判者缺乏其所需的证据,以及一个已配置但未应用的模型设置。

英文摘要

We present a framework for evaluating and improving a large-scale, multi-agent shopping assistant in production, and report lessons from its use. Offline evaluation of such a system faces three obstacles. (i) A logged conversation cannot be replayed against a modified system, because a different response changes every turn that follows. (ii) The unchanged system itself varies from run to run. Its LLM components are stochastic, and in product search the available products, their prices, and the customer's personalization signals change. (iii) Aggregate quality scores combine distinct behaviors, so they show that quality has changed but not which behavior caused the change. Our framework addresses each obstacle in turn. For a reported behavior, an Evaluation Harness generates targeted assertions and a fixed cohort of customer scenarios. It then reproduces the behavior in a local instance of the assistant through grounded user simulation. Instead of replaying the log, the simulator writes new customer turns conditioned on the recorded messages and context. Repeated runs of the unchanged system form a stored baseline. An Improvement Orchestrator turns the assertion results into hypotheses, implements each as an isolated modification, and compares it with the baseline using paired percentile bootstrap intervals over scenario-level differences. When an investigation ends, the harness may propose revisions to future evaluations, subject to human approval and without altering past decisions. We report production investigations with this framework. Assertion profiles showed which positions of a product carousel a failure affected, and repeated runs distinguished a real improvement from run-to-run fluctuation. Audits of the evaluation itself found a judge that lacked the evidence it needed and a model setting that was configured but not applied.

Comments21 pages, 2 figures, 1 table

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑