发表机构
Beijing University of Posts and Telecommunications; Singapore University of Technology and Design(北京邮电大学; 新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对异构大语言模型实体代理协作面临的通信开销大、协调质量受LLMs能力限制和行动延迟等挑战,提出基于轻量级数字孪生的LDT-Coord框架,通过特定方法减少通信开销,实现与传统方法相当的任务成功率且保持鲁棒性。
AI 中文摘要
由异构大语言模型驱动的实体代理团队正广泛应用于智能工厂、仓库和服务机器人等物理人工智能领域。为实现此类代理团队的协作,需要在有限网络资源下可靠运行的高效协调机制。现有基于多轮自然语言对话的异构大语言模型-代理协调框架存在三个相互关联的挑战。本文提出了LDT-Coord,一个基于轻量级数字孪生构建的网络协调框架。具体而言,各代理独立选择预期行动并向数字孪生服务器报告行动决策和对共享资源的结构化时间约束,数字孪生执行无训练、基于规则的编排算法解决跨代理冲突并返回协调指令。为进一步减少通信开销,将代理报告控制制定为约束部分可观测马尔可夫决策过程并用PPO-拉格朗日算法求解。仿真结果表明,LDT-Coord在降低通信开销超70倍并在大语言模型异构性下保持鲁棒性的同时,实现了与传统协调方法相当的任务成功率。
英文摘要
Embodied agent teams powered by heterogeneous large language models (LLMs) are being widely deployed in physical artificial intelligence such as smart factories, warehouses, and service robotics. To enable collaboration among such an agent team, efficient coordination mechanisms that operate reliably under limited network resources are required. However, existing heterogeneous LLM-agent coordination frameworks that rely on multi-round natural-language-based conversations introduce three coupled challenges. First, inter-agent dialogue incurs communication overhead that grows rapidly with team size. Second, the quality of coordination is constrained by the heterogeneous capabilities of the agent team's LLMs. Third, agents may suffer from action delays due to iterative negotiation. To address these challenges, we propose LDT-Coord, a networked coordination framework built upon a lightweight digital twin (DT). Specifically, each agent independently selects its intended action and reports both the action decision and a structured temporal constraint over shared resources to the DT server, thereby decoupling coordination performance from natural-language reasoning ability. Then, DT executes a training-free, rule-based orchestrator algorithm to resolve cross-agent conflicts and returns coordination instructions to prevent such conflicts. To further reduce communication overhead, we formulate agent reporting control as a constrained partially observable Markov decision process (C-POMDP) and solve it with the PPO-Lagrangian algorithm. Simulation results show that LDT-Coord achieves a task success rate comparable to conventional coordination methods while reducing communication overhead by more than 70x and maintaining robustness under LLM heterogeneity.
Comments14 pages, 6 figures, 5 tables