AI 中文总结
针对多智能体规划中隐私受限合作难题,提出RELIC框架,通过揭示原则学习可解释可组合技能,各智能体私用大语言模型引导搜索优化技能,协调器依团队性能评估更新,实现跨智能体转移,引入隐私保护技能学习新范式。
AI 中文摘要
当智能体必须在保持内部实现私密的同时提高专业决策技能时,多智能体规划变得更加困难。这种情况出现在智能体独立开发、暴露不同接口和能力且必须在不共享可执行策略的情况下进行协调时。先前研究大多假设集中优化、共享策略访问或通用技能表示,因此不太适合隐私受限的合作。我们引入了RELIC,一个通过揭示原则学习可解释和可组合技能的框架。每个智能体通过私有的大语言模型引导搜索来优化自己的程序性技能,而一个可信的协调器仅通过团队级性能评估提议的更新。成功行为不作为代码广播,而是被抽象为可移植的原则,其他智能体可以在自己的接口中实例化并与本地策略重组。这将协调与实现共享分开,实现了异构技能签名下的跨智能体转移。RELIC因此为多智能体规划中隐私保护技能学习和协调引入了一种新范式。
英文摘要
Multi-agent planning becomes substantially harder when agents must improve specialized decision-making skills while keeping their executable implementations private. This setting arises when independently developed agents expose heterogeneous interfaces, observations, and capabilities, yet must coordinate under a shared team objective. Existing approaches commonly rely on centralized optimization, shared policy access, or common skill representations, assumptions that limit knowledge reuse when function signatures differ. We introduce RELIC, a framework for learning interpretable and composable programmatic skills through revealed principles. Each agent improves its own executable skill locally, while useful decision logic and coordination patterns are distilled into compact textual principles. Rather than requiring direct program exchange, these abstractions can be re-instantiated under agent-specific interfaces and reused across incompatible implementation spaces. A shared principle memory accumulates transferable knowledge and promotes abstractions that repeatedly improve team-level performance. This separation allows discoveries made by one agent to guide others while preserving local executable implementations and decentralized execution. RELIC therefore supports strategic transfer across both heterogeneous-role and shared-role cooperative teams. Extensive experiments across routing, scheduling, combinatorial optimization, and distributed coordination settings demonstrate RELIC's effectiveness against independent and joint LLM-based search methods, together with consistent benefits across task structures and LLM backbones.
Commentsv1 accepted at LM4Plan Workshop @ ICML 2026; v2 is the full paper version; Kiet, Pham, and Chinh contributed equally in v2