arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11207cs.AI

用于实现协作式对话结果的多大型语言模型智能体系统的动态管控

Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes

Alexander Liss, Nicholas Desmond, Santiago Gil Gallego

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对多LLM智能体对话崩溃问题,提出体验编排器EO管控层,通过三种机制提升顾问联系率,在6万次模拟中取得显著效果,为多智能体协作提供新方案。

中文摘要 AI 辅助

当两个具有结构对立目标的大型语言模型(LLM)智能体进行多轮交互时,由于缺乏共享目标函数,不会产生竞争,反而会导致对话崩溃:访客智能体屈服,站点智能体停止调整策略,对话终止且未达成任何一方的既定目标。本文探究控制理论管控层是否可替代缺失的目标函数。体验编排器(Experience Orchestrator, EO)在模拟金融服务环境中解决该问题,其中站点智能体引导访客联系顾问,而访客保持符合心理现实的抗拒状态。EO通过三种机制管控联合轨迹:上下文多臂赌博机(Contextual Bandit, CB),其选择的内容臂基于真实网络分析校准;PID控制器,通过动态模式约束强制行为一致性;部分可观察马尔可夫决策过程(POMDP)信念跟踪器,维护访客意图的概率模型。在6万次模拟中,EO使高意图顾问联系率提升了32个百分点(对比朴素LLM对照组,EO的高意图顾问联系率为78.1%,对照组为46.1%),且CB变体选择解释了97%的因素间结果方差,证实管控策略而非环境初始条件决定轨迹最终走向。角色层面分析揭示两种不同场景:对于无自然转化倾向的访客,管控层是系统能否正常运行的关键;对于已接近达成一致的访客,朴素LLM的共情默认设置基本足够。所有发现均基于LLM间模拟,PID控制器未针对真实人类的不可预测性进行校准,在真实流量上验证EO是关键的下一步。

英文摘要

When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective. This paper asks whether a control-theoretic governance layer can substitute for that missing goal function. The Experience Orchestrator (EO) addresses this in a simulated financial services environment where a site agent guides a visitor toward advisor contact while the visitor maintains psychologically realistic resistance. EO governs the joint trajectory through three mechanisms: a Contextual Bandit (CB) that selects content arms calibrated from real-world web analytics, a PID controller that enforces behavioral consistency via dynamic schema constraints, and a POMDP belief tracker that maintains a probabilistic model of visitor intent. Across 60,000 simulations, EO achieves a +32 percentage point lift in high-intent advisor contact rate (78.1% vs. 46.1% over a naive LLM control), with CB variant selection accounting for 97% of between-factor outcome variance -- confirming that the governance policy, not environmental initial conditions, determines where trajectories end up. Persona-level analysis reveals two distinct regimes: for visitors with no natural inclination toward conversion, the governance layer is the difference between a functional system and a non-functional one; for visitors already near alignment, a naive LLM's empathetic defaults are largely sufficient. All findings are conditional on LLM-to-LLM simulation. The PID controller has not been calibrated against real human unpredictability, and validating EO on live traffic is the critical next step.

补充信息

↑