arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05758cs.AI

从单体混合到智能体编排:大规模对话助手的动态响应

From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale

Cen Mia Zhao, Peng Wang, Chuan Shi, Yufeng Zhang, Ying Lyu, Wanmeng Ren, Robert Xue, Claire Na Cheng, Yashar Mehdad

首次发表
浏览论文内容

中文总结 AI 辅助

针对大规模对话助手,提出从单体混合到智能体编排的动态响应方法,通过类型化工具和验证上下文降低幻觉与升级率,并显著优化延迟与成本。

中文摘要 AI 辅助

对话助手可以在单一模型路径中混合检索、动作选择、升级和措辞,或者将这些角色分离。我们报告了在一个大型住宿市场(每月数百万次对话,11种语言,P90延迟10秒)中客户支持助手的生产迁移。动态响应(DR)用有界ReAct编排器(基于类型化工具)加上一个较小的生成器(根据后端验证的上下文契约进行写作)取代了单一的Qwen3-235B-A22B混合响应器。由于迁移还改变了提示、对齐和服务,我们将每个效果归因于其原因,并仅将在相同重放回合上测量的效果视为架构效果:类型化实体选择将预订选择器移至精确率优先的工作点(精确率从8.3%提升至89.1%,召回率从75.2%降至67.3%),类型化动作ID与成员检查消除了观察到的结构化动作幻觉(从2.14%降至0.0%)。低梯度A/B测试重现了重放升级减少:硬升级响应从5.60%降至3.08%,软升级响应从9.56%降至2.49%,而生产转交手量大致保持稳定;自解决呈方向性(+5.1个百分点,95%置信区间[-2, +12])。服务优化将编排器P90延迟从3.87秒降至2.24秒,GPU占用减少约三分之一,自托管将估计的年度模型服务成本降低了一个数量级以上。

英文摘要

Conversational assistants can blend retrieval, action selection, escalation, and wording in a single model path, or separate those roles. We report a production migration of a customer-support assistant at a large accommodation marketplace (millions of conversations per month, 11 languages, 10-second P90). Dynamic Response (DR) replaces a single Qwen3-235B-A22B blended responder with a bounded ReAct orchestrator over typed tools plus a smaller generator that writes from a backend-validated context contract. Because the migration also changed prompts, alignment, and serving, we attribute each effect to its cause and claim as architecture effects only those measured on identical replayed turns: typed entity selection moves the reservation selector to a precision-first operating point (precision 8.3% to 89.1%, recall 75.2% to 67.3%), and typed action IDs with a membership check remove observed structured-action hallucination (2.14% to 0.0%). A low-ramp A/B test reproduces the replay escalation reductions: hard-escalation responses fall from 5.60% to 3.08% and soft-escalation responses from 9.56% to 2.49%, while production handoff volume holds roughly steady; self-solve is directional (+5.1 points, 95% CI [-2, +12]). Serving optimizations cut orchestrator P90 latency from 3.87s to 2.24s on a GPU footprint reduced by roughly one-third, and self-hosting reduces estimated annual model-serving cost by more than an order of magnitude.

发表机构

  • Airbnb, Inc.(爱彼迎公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑