arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2608.11658cs.LGcs.AIcs.MA

智能体策略组合是否安全?重新思考合作多智能体强化学习中的后继特征迁移

Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning

  • The Hong Kong University of Science and Technology(香港科技大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

Zijian Zhao, Sen Li

AI总结:

针对多智能体强化学习中独立策略组合不安全的问题,提出MA-USFA分层方法,兼顾策略组合的安全性与灵活性,无需任务适配即可部署。

AI中文摘要:

许多强化学习系统,从车队管理到交通信号控制,必须满足部署后动态变化的目标,为每个新目标重新训练策略的成本高得令人望而却步。对于单智能体,该问题已得到充分理解:结合广义策略改进的后继特征及其通用扩展,可将已学习策略库重新组合为适用于任何新目标的策略,且保证结果绝不劣于库中的任何策略。然而,多智能体迁移受到的关注少得多,让每个智能体独立重组自身库的常见做法虽沿用了该思路,却未继承其保证。我们证明,这种独立组合会产生比库中所有策略都差的联合行为,因为重组队友会改变每个智能体面临的环境,使其依赖的价值失效,这是单智能体不存在的失败。我们进一步表明,唯一无条件安全的固定规则是同步组合,它将整个团队转移到一个联合训练的策略,但无法服务于为不同智能体分配不同目标的情况。为同时实现安全性和灵活性,我们提出MA-USFA,这是一种两层分层方法:下层是通用后继特征近似器,在每个智能体的队友目标条件下预测其后继特征;上层组合器在智能体间选择每个智能体应遵循的库条目,并提供单智能体价值无法表示的跨智能体修正。该方法在目标分布上训练一次,部署时无需针对每个任务进行调整。

英文摘要:

Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for each new objective is prohibitively expensive. For a single agent, this problem is well understood: successor features with generalized policy improvement, together with their universal extension, recombine a library of learned policies into a policy for any new objective, with a guarantee that the result is never worse than any policy in the library. However, multi-agent transfer has received far less attention, and the common practice of letting each agent recombine its own library independently inherits the recipe but not the guarantee. We prove that this independent composition can produce joint behavior strictly worse than every policy in the library, because recombining teammates changes the environment each agent faces and invalidates the values it relies on, a failure with no single-agent counterpart. We further show that the only unconditionally safe fixed rule is synchronized composition, which moves the whole team to one jointly trained policy but cannot serve objectives that assign different goals to different agents. To attain safety and flexibility at once, we propose MA-USFA, a hierarchical method with two layers: a lower layer of universal successor feature approximators that predicts each agent's successor features while conditioned on its teammates' objectives, and an upper composer that selects, across agents, which library entry each agent should follow and supplies the cross-agent correction a per-agent value cannot represent. Trained once over the distribution of objectives, it is applied at deployment with no per-task adaptation.

↑