arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12321cs.MAcs.GTcs.LG

公共物品困境中多智能体学习产生的空间模式形成

Spatial Pattern Formation from Multi-Agent Learning in Public Goods Dilemmas

Yefei Zhang, Yuxuan Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探究公共物品困境中多智能体学习的空间模式形成,发现学习率影响集体福利,特定学习率组合致最大福利损失,收取拥挤成本可恢复部分损失,模式反映训练历史而非渐近结果。

中文摘要 AI 辅助

空间公共物品模型表明,向更丰富位置的规定性移动可产生空间模式。本研究探究当智能体学习移动位置时此类模式如何出现,以及学习率如何影响其对集体福利的后果。固定数量的合作者与背叛者群体使用表格型Q学习算法,基于局部观测独立学习移动策略。合作者学习会在资源峰值周围形成集群,而共同适应会改变这些集群的强度与运动。在固定训练预算下,当合作者以高学习率学习、背叛者以低学习率学习时,会出现最大的福利损失。在该机制的部分区域,学习到的策略还会产生由共享方向偏好支撑的移动带。支撑移动的条件会随进一步训练而变化,因此这些模式反映的是训练历史,而非既定的渐近结果。在测试的合作者学习的学习率条件下,平均集体福利低于随机移动,因为拥挤加剧超过了资源收益的增加。在学习期间向智能体收取其对他人造成的拥挤成本,可在测试条件下恢复大部分福利损失。这些结果将学习率与个体奖励驱动的空间组织的出现及福利成本关联起来。

英文摘要

Spatial public goods models show that prescribed movement toward richer locations can generate spatial patterns. We ask how such patterns emerge when agents learn where to move and how learning rates shape their consequences for collective welfare. Fixed populations of cooperators and defectors independently learn movement policies using tabular Q-learning and local observations. Cooperator learning generates clusters around resource peaks, while co-adaptation changes their strength and motion. At a fixed training budget, the largest welfare losses occur when cooperators learn at high rates and defectors at low rates. In part of this regime, learned policies also generate traveling bands supported by a shared directional preference. The conditions supporting travel change with further training, so these patterns reflect training history rather than an established asymptotic outcome. Across the tested learning-rate conditions with cooperator learning, mean collective welfare falls below random movement because increased crowding outweighs gains in resource benefit. Charging agents for the crowding they impose on others during learning recovers much of the welfare loss in the tested conditions. These results connect learning rates to the emergence and welfare costs of spatial organization driven by individual rewards.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑