基于收益的高阶复制子动态的稳定极限
Stabilization Limits of Payoff-Based Higher-Order Replicator Dynamics
浏览论文内容
中文总结 AI 辅助
该研究分析基于收益的高阶复制子动态的稳定极限,证明辅助线性时不变系统非无源时对应博弈内部纳什均衡不稳定,特定博弈无法被指定高阶RD局部稳定,放松纳什平稳性后广义指数RD可稳定熵正则化近似纳什均衡。
中文摘要 AI 辅助
复制子动态(Replicator Dynamics,RD)是博弈学习领域的基础模型,连接了演化博弈理论与在线学习。本文研究基于收益的高阶RD变体,其可表示为积分器与辅助线性时不变(Linear Time-Invariant,LTI)系统并联后与softmax映射的级联互联结构。我们分析该纳什平稳学习规则下纳什均衡的可学习性:首先,回顾近期结果——当辅助LTI系统严格无源时,该动态收敛至纳什均衡;并证明逆无源结果:若辅助LTI系统非无源,则存在静态严格收缩博弈,其内部纳什均衡在闭环学习动态下不稳定。其次,证明存在一类具有孤立内部纳什均衡的博弈,无法被任何辅助LTI系统渐近稳定且严格正则的基于收益的高阶RD局部渐近稳定。最后,证明若放松纳什平稳性(即所有纳什均衡均为学习动态的平稳点),则广义指数RD(Generalized Exponential RD,Ex-RD)可对任意连续可微博弈局部渐近稳定logit均衡,该稳定均衡可视为熵正则化近似纳什均衡。
英文摘要
Replicator dynamics (RD) is a fundamental model in learning in games, connecting evolutionary game theory and online learning. This paper studies payoff-based higher-order variants of RD represented as a cascade interconnection between an integrator in parallel with an auxiliary linear time-invariant (LTI) system and the softmax mapping. We investigate learnability of Nash equilibria under this Nash-stationary learning rule. First, we revisit recent results that establish convergence to Nash Equilibrium whenever the auxiliary LTI system is strictly passive and prove a converse passivity result: if the auxiliary LTI system is not passive, then there exists a static strictly contractive game whose interior Nash equilibrium is unstable under the closed-loop learning dynamics. Second, we show that there exists a class of games with isolated interior Nash equilibria that cannot be locally asymptotically stabilized by any payoff-based higher-order RD whose auxiliary LTI system is asymptotically stable and strictly proper. Finally, we show that if Nash stationarity (i.e., all Nash equilibria are stationary points of the learning dynamics) is relaxed, then generalized exponential RD (Ex-RD) can locally asymptotically stabilize a logit equilibrium for any continuously differentiable game. The stabilized equilibrium can be viewed as an entropy-regularized approximate Nash equilibrium.