发表机构
Centific(森蒂菲克(Centific))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对企业对话智能体的标注瓶颈,提出RL-ADA协同进化框架,用世界反馈取代人工标注,在银行场景中消除路由错误并提升PASS率,还发现上下文伪装的对抗策略。
AI 中文摘要
在企业客户支持场景中部署任务型对话智能体面临持续的标注瓶颈:鲁棒训练需要大规模带标注的交互数据,但企业对话日志具有隐私敏感性且标注成本高昂,同时用户行为的演化速度远超标注流程的跟进速度。我们提出RL-ADA(带对抗性对话智能体的强化学习),一种协同进化训练框架,通过用「世界反馈」取代人工标注来消除该瓶颈:直接源自可测量交互结果的基于后果的奖励信号。客户支持智能体(DA,30亿参数)与对抗性客户智能体(CA,70亿参数)在由固定自动评判器引导的对抗场景中协同进化:DA因正确处理多轮客户对话并达成成功解决而获得奖励,CA因生成隐藏意图的逼真话语以造成错误路由而获得奖励,通过相互对立但独立构建的奖励形成不对称对抗压力。隔离式训练环境会基于先前失败的对话文本迭代重新训练较弱的智能体,全程无需人工标注。在银行客户支持的概念验证中,仅通过自动场景奖励(无标注数据),工具路由错误被消除,严格的端到端PASS率在五次协同进化周期中翻倍。我们还观察到「上下文伪装」的出现:CA仅通过奖励压力学习将意图嵌入密集逼真的客户细节中的对抗策略,对企业红队测试和鲁棒性评估具有直接意义。
英文摘要
Deploying task-oriented dialogue agents in enterprise customer support faces a persistent annotation bottleneck: robust training requires labelled interaction data at scale, yet enterprise conversational logs are privacy-sensitive and expensive to annotate, while user behaviour evolves faster than labelling pipelines can keep pace. We present RL-ADA (Reinforcement Learning with Adversarial Dialogue Agents), a co-evolutionary training framework that eliminates this bottleneck by replacing human labels with \emph{world feedback}: consequence-based reward signals derived directly from measurable interaction outcomes. A Customer Support Agent (DA, 3B parameters) and an Adversarial Customer Agent (CA, 7B parameters) co-evolve in an adversarial arena guided by a fixed automated judge: the DA is rewarded for correctly handling multi-turn customer conversations to successful resolution, while the CA is rewarded for producing realistic, intent-concealing utterances that cause misroutes, creating asymmetric adversarial pressure through opposing but independently structured rewards. An isolation gym iteratively retrains the weaker agent on prior-failure transcripts, requiring no human annotation at any stage. In a banking customer support proof of concept, tool-routing errors are eliminated and the strict end-to-end PASS rate doubles over five co-evolutionary cycles, driven solely by automated arena reward with no labelled data. We additionally observe the emergence of \textbf{Contextual Camouflage}, an adversarial strategy in which the CA learns to embed intent within dense realistic customer detail purely from reward pressure, with direct implications for enterprise red-teaming and robustness evaluation.