arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01630cs.MA

合作学习之后:多智能体强化学习中的梯度路由与优化器依赖的维持

After Cooperation Is Learned: Gradient Routing and Optimizer-Dependent Maintenance in Multi-Agent Reinforcement Learning

Chaoyuan Hao, Wentao Yue, Tianyou Lai, Hongji Li, Jiayi Zhou, Qingyu Mao, Qilei Li

AI总结:

本研究通过右删失事件时间框架和梯度路由控制实验,发现直接价值梯度访问在特定缩放条件下增加合作维持风险,而非评论家本身的普遍问题。

AI中文摘要:

合作式多智能体强化学习通常通过从随机初始化中发现合作来评估,但持续优化是否会破坏已学习的合作仍是一个未解问题。演员-评论家比较也可能将评论家的存在与进入共享演员表征的价值梯度混为一谈。我们研究合作维持,定义为行为上验证的合作策略在持续训练下存活的能力。我们将维持问题形式化为右删失事件时间问题,并比较匹配的热启动:X0允许价值损失梯度更新共享演员特征,X1保留评论家但阻断这些梯度,X5移除学习的评论家作为无评论家参考。这隔离了直接价值梯度访问,同时控制了初始化、评论家计算和评估。正奖励缩放保留策略偏好和均衡,同时扰动学习动态。梯度审计确认了预期的路由路径,冻结策略躯干扰动探测路由诱导的更新是否与局部合作边界对齐。在确证性MinEx和CleanUp-lite实验中,较高缩放比例选择性地增加X0中的维持敏感性;X1保持接近删失上限,X5在测试设置中无确认事件。在CleanUp-lite中,路由乘以缩放位移与局部合作边际减少相关;MinEx显示较弱且依赖优化器的效应。这些结果识别出与直接价值梯度路由相关的条件性、缩放敏感的维持风险,而非评论家的普遍失败。

英文摘要:

Cooperative MARL is commonly evaluated through cooperation discovery from random initialization, leaving open whether continued optimization can destabilize learned cooperation. Actor-critic comparisons can also conflate critic presence with value gradients entering shared actor representations. We study cooperation maintenance, defined as the survival of a behaviorally verified cooperative policy under continued training. We formulate maintenance as a right-censored event-time problem and compare matched warm starts: X0 allows value loss gradients to update shared actor features, X1 retains the critic while blocking those gradients, and X5 removes the learned critic as a critic-free reference. This isolates direct value-gradient access while controlling initialization, critic computation, and evaluation. Positive reward scaling preserves strategic preferences and equilibria while perturbing learning dynamics. Gradient audits confirm the intended routing pathways, and frozen-policy torso perturbations probe whether route-induced updates align with local cooperation boundaries. In confirmatory MinEx and CleanUp-lite experiments, higher scales selectively increase maintenance sensitivity in X0; X1 remains near the censoring ceiling, and X5 has no confirmed events in the tested settings. In CleanUp-lite, route-by-scale displacement is associated with reduced local cooperation margins; MinEx shows a weaker, optimizer-dependent effect. These results identify a conditional, scale-sensitive maintenance risk associated with direct value-gradient routing rather than a universal failure of critics.

↑