arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循证优化:面向客户服务语言模型代理的跟踪驱动优化

Evidence-in-the-Loop: Trace-Driven Optimization for Customer-Service LLM Agents

Chunming Wu, Dafei Qiu, Congde Yuan, Charles Quan, Jun Wu, Suipeng Li, Mo Wu, Gavin Xie, Hope Chen, Max Yao

arXiv 2607.18039首次发表:更新:

AI 中文总结

研究面向客户服务语言模型代理的优化,提出循证客户服务代理工作流程,通过多种技术构建证据,结合政策引导编排。贡献混合RAG证据构建、循证问题/行动决策、跟踪驱动的RAG和重排器改进三种部署模式。

AI 中文摘要

生产环境中的客户服务机器人必须在迭代发布中提高回答质量,同时大语言模型不能突破证据边界、政策规则或人工交接保障。我们展示了一种在实际客户服务环境中部署的循证客户服务代理工作流程。BM25召回、问题标题向量召回、问题描述向量召回、加权RRF融合和交叉编码器重排为受控制的大语言模型决策构建有依据的常见问题解答证据。然后,政策引导编排将此RAG证据与特定场景规则证据、对话记忆以及固定LangGraph DAG中的澄清状态相结合。本文贡献了三种可复用的部署模式:混合RAG证据构建,多通道检索和重排产生可审计的常见问题解答候选;循证问题/行动决策,循证决策模块从类型化常见问题解答证据和特定场景规则证据中选择问题/行动;跟踪驱动的RAG和重排器改进,跟踪诊断故障是来自召回、排序、最终候选选择、澄清规则衍生证据还是行动政策,重排器微调不仅评估领域内收益,还评估遗忘风险。

英文摘要

Production customer-service bots must improve answer quality across iterative releases, yet large language models must not bypass evidence boundaries, policy rules, or human-handoff safeguards. We present an \textbf{Evidence-Grounded Customer-Service Agent Workflow} deployed in a real-world customer-service setting. BM25 recall, issue-title-vector recall, issue-description-vector recall, weighted RRF fusion, and cross-encoder reranking construct grounded FAQ evidence for controlled LLM decisions. Policy-guided orchestration then combines this RAG evidence with scenario-specific rule evidence, conversation memory, and clarification state inside a fixed LangGraph DAG~\cite{langgraph2024}. The paper contributes three reusable deployment patterns: \textbf{hybrid RAG evidence construction}, where multi-channel retrieval and reranking produce auditable FAQ candidates; \textbf{evidence-grounded issue/action decision}, where an Evidence-Grounded Decision Module selects an issue/action from typed FAQ evidence and scenario-specific rule evidence; and \textbf{trace-driven RAG and reranker improvement}, where traces diagnose whether failures come from recall, ranking, final candidate selection, clarification, rule-derived evidence, or action policy, and where reranker fine-tuning is evaluated not only for in-domain gain but also for forgetting risk.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑