发表机构
The University of Texas at Austin(德克萨斯大学奥斯汀分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对离线多智能体强化学习中生成策略与价值优化的权衡,提出SCOUT框架,通过测试时动作细化结合流匹配先验与分解值函数,利用Stein变分梯度下降实现可扩展协调,在基准上取得最优性能。
AI 中文摘要
离线多智能体强化学习(MARL)面临一个持续的权衡。具有表达力的生成式策略能够表示数据中的多模态协调,但无法区分高价值区域,而价值优化策略利用学习到的Q函数,但将多模态坍缩为单一主导模态。单个智能体的模态坍缩可能破坏联合协调,而跨智能体的同时漂移可能将联合策略推向动作空间的未知区域。我们提出通过最优统一传输的可扩展协调(SCOUT),这是第一个将生成式基础模型与学习到的价值函数通过测试时动作细化相结合的离线MARL框架。SCOUT训练两个解耦组件:一个流匹配行为先验和一个分解价值函数。在测试时,它通过Stein变分梯度下降将行为样本传输到高价值区域。传输步数控制自适应的测试时缩放,取代固定的正则化系数。在个体-全局-最大(IGM)原则下,我们证明了联合软价值差距上的单项KL界,该界随着传输收敛而消失,并有一个与IGM违反成比例的不可约加性残差。实验上,SCOUT在离散和连续离线MARL基准上取得了最佳平均性能,并在所有离线到在线配置中产生了性能改进。
英文摘要
Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time action refinement. SCOUT trains two decoupled components: a flow-matching behavioral prior and a decomposed value function. At test-time, it transports behavioral samples toward high-value regions via Stein variational gradient descent. The number of transport steps controls adaptive test-time scaling, replacing a fixed regularization coefficient. Under the individual-global-max (IGM) principle, we prove a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation. Empirically, SCOUT achieves the best average performance across discrete and continuous offline MARL benchmarks and yields performance improvements in all offline-to-online configurations.
CommentsNeurIPS 2026