arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体在线学习中偏好能够预测和无法预测的内容

What preferences can - and cannot - predict in multi-agent online learning

Omar Abbadi, Rida Laraki, Panayotis Mertikopoulos

arXiv 2608.13810首次发表:更新:

发表机构

Univ. Mohammed VI Polytechnic; Univ. Grenoble Alpes; CNRS; Inria; Grenoble INP; LIG(穆罕默德六世理工大学; 格勒诺布尔大学; 法国国家科学研究中心; 法国国家信息与自动化研究所; 格勒诺布尔国立高等工业学院; 信息与信号处理实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究博弈中偏好与无遗憾学习动态的关联,证明子博弈中偏好可刻画渐近稳定性,但一般博弈中偏好不足以决定动态稳定性,进而提出聚合偏离下的弹性概念保证纯策略跨度的渐近稳定性。

AI 中文摘要

本文研究博弈中基于序数偏好的解概念与博弈动态长期行为之间的相互作用,尤其关注博弈的组合数据(即偏好图)在多大程度上决定无遗憾学习动态(如正则化领导者跟随算法FTRL)的结果。一方面,我们证明每个动态稳定集的骨架(即其包含的纯策略配置集)也必须是偏好稳定的,即该集必须对有利可图的偏离是封闭的。随后我们提出反向问题:偏好何时能决定玩家学习动态的长期行为?我们首先证明,在子博弈(即通过限制玩家动作集得到的纯策略配置子集)的情况下,偏好可刻画渐近稳定性。然而超出该情况后,动态稳定性与偏好稳定性的等价性不再成立:具体而言,我们构造了一个三人博弈,其中存在一个偏好稳定集,其跨度是动态不稳定的,由此表明偏好不足以作为动态稳定性的判据。接着我们通过聚合偏离下的弹性概念弥合这一差距,该概念是一种易于验证的基于收益的条件,可保证任意纯策略跨度的渐近稳定性。

英文摘要

We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game -- its preference graph -- determine the outcomes of no-regret learning dynamics -- such as follow-the-regularized-leader (FTRL). In one direction, we show that the skeleton of every dynamically stable set (i.e. the set of pure profiles it contains) must also be preferentially stable, that is, it must be closed under profitable deviations. We then ask the converse question: when do preferences determine the long-run behavior of the players' learning dynamics? We begin by showing that preferences characterize asymptotic stability in the case of subgames -- i.e. subsets of pure profiles obtained by restricting players' action sets. Beyond this case however, the equivalence between dynamic and preferential stability collapses: concretely, we construct a three-player game with a preferentially stable set whose span is dynamically unstable, showing in this way that preferences do not suffice as a criterion of dynamic stability. We then bridge this gap via the notion of resilience under aggregate deviations, an easy-to-check payoff-based condition that guarantees asymptotic stability of arbitrary spans of pure strategies.

Comments56 pages, 15 figures; oral presentation at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑