一种具有均值场环境反馈的多智能体Q学习的进化计算框架
An Evolutionary Computation Framework for Multi-Agent Q-Learning with Mean-Field Environmental Feedback
浏览论文内容
中文总结 AI 辅助
提出一种耦合学习-环境的均值场框架,用于分析网络化多智能体Q学习与环境反馈的相互作用,并通过蒙特卡洛模拟验证其能再现宏观合作与环境轨迹。
中文摘要 AI 辅助
网络化群体中的多智能体强化学习受个体适应、局部遭遇和变化的环境条件之间相互作用的影响。为了研究这种相互作用,我们构建了一个耦合的学习-环境模型,其中智能体在固定图上更新无状态的$Q$值,而其群体平均行为驱动一个环境变量,该变量动态修改收益矩阵。在一阶均值场闭合下,我们推导出$Q$值群体分布的确定性输运方程,并将其与环境状态的投影离散更新相耦合。所得模型在随机正则图、Erdős–Rényi图、Barabási–Albert图和随机几何图上与有限网络蒙特卡洛模拟进行了评估。在所测试的参数范围内,均值场系统再现了主要的宏观合作和环境轨迹,且轨迹级别的均方根误差通常随群体规模和平均度数的增加而减小。分析进一步表明,环境反馈重塑了学习到的动作值排序,而强化反馈可产生对初始学习偏差和资源水平的显著依赖。环境时间尺度也起着重要作用:快速响应可在学习适应之前将资源状态驱动到边界,而较慢的响应则保留行为学习与环境恢复之间的相互作用。这些结果为耦合强化学习和环境动力学提供了群体层面的描述,并刻画了在所测试网络和参数范围内均值场近似的性能。
英文摘要
Multi-agent reinforcement learning in networked populations is governed by the interaction between individual adaptation, local encounters, and changing environmental conditions. To study this interaction, we formulate a coupled learning--environment model in which agents update stateless $Q$-values on a fixed graph, while their population-average behavior drives an environmental variable that dynamically modifies the payoff matrix. Under a first-order mean-field closure, we derive a deterministic transport equation for the population distribution of $Q$-values and couple it with a projected discrete update for the environmental state. The resulting model is evaluated against finite-network Monte Carlo simulations on random regular, Erdős--Rényi, Barabási--Albert, and random geometric graphs. Across the tested parameter ranges, the mean-field system reproduces the main macroscopic cooperation and environmental trajectories, and the trajectory-level root-mean-square error generally decreases with population size and average degree. The analysis further shows that environmental feedback reshapes the learned action-value ordering, while reinforcing feedback can produce pronounced dependence on the initial learning bias and resource level. The environmental timescale also plays an important role: a rapid response can drive the resource state to a boundary before learning adapts, whereas a slower response preserves the interaction between behavioral learning and environmental recovery. These results provide a population-level description of coupled reinforcement learning and environmental dynamics and characterize the performance of the mean-field approximation within the tested network and parameter ranges.
发表机构
- Northwest A & F University(西北农林科技大学)
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。