arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

G2MAF:多智能体流策略的测试时梯度引导

G2MAF: Test-Time Gradient Guidance for Multi-Agent Flow Policies

Guowei Zou, Haitao Wang, Guoxin Wang, Zhiquan Chen, Beiwen Zhang, Guojie Wang, Hejun Wu

arXiv 2609.31286首次发表:更新:

AI 中文总结

提出G2MAF框架,利用全局归一化投影评论家梯度在测试时引导多智能体流策略修正,在24个MPE和SMAC设置中改进20个,平均增益约9%,延迟仅增6%。

AI 中文摘要

离线多智能体强化学习(MARL)从固定数据集中学习协作策略,无需进一步的环境交互,且学习到的策略在部署时被冻结。这种冻结策略通常提出单一联合动作并在部署时直接执行。然而,这种一次性部署往往承诺一个次优的提议,即使更好的邻近替代方案与行为数据保持一致。为解决此问题,我们提出梯度引导多智能体流(G2MAF),一种在测试时优化联合策略的细化框架。G2MAF应用一个全局归一化、投影的评论家梯度来引导和协调所有智能体的修正,同时保持动作既可行又接近冻结策略的提议。在24个MPE和SMAC设置中,其规范变体改进了20个冻结设置,在MPE上平均相对增益为9.2%,在SMAC上为8.9%,模型推理延迟仅增加约6%。

英文摘要

Offline multi-agent reinforcement learning (MARL) learns cooperative policies from fixed datasets without further environment interaction and a learned policy is frozen at deployment. Such a frozen policy typically proposes a single joint action and executes it directly at deployment time. However, this one-shot deployment often commits to a suboptimal proposal, even when better nearby alternatives remain consistent with the behavior data. To address this issue, we propose Gradient Guided Multi Agent Flow (G2MAF), a refinement framework for optimizing joint policies at test-time. G2MAF applies one globally normalized, projected critic gradient to guide and coordinate all agents' corrections while keeping the action both feasible and close to the frozen policy proposal. Across 24 MPE and SMAC settings, its canonical variant improves 20 frozen settings, with mean relative gains of 9.2% on MPE and 8.9% on SMAC, with model inference latency increased by about 6% only.

Comments24 pages, including appendices. Project page: https://g2maf.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑