发表机构
The University of Osaka; University of Glasgow; Université Polytechnique Hauts-de-France; INSA Hauts-de-France; RIKEN Center for Advanced Intelligence Project(大阪大学; 格拉斯哥大学; 上法兰西理工学院; 上法兰西国立应用科学学院; 理化学研究所先进智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过霍奇分解揭示势博弈替代优先级协调的遗漏,证明其调和分量不可消除,并给出线性收益下的闭式误差下界与设计极限。
AI 中文摘要
在去中心化的优先级协调中,智能体宣布优先级水平,共享资源按降序为其服务,如同在无信号交叉路口的情况;这些水平构成了分层控制器的决策层。此类交互在分析和设计中通常被一个势博弈(即一个共同目标)所替代。本文确定了这种替代所遗漏的内容,利用激励的霍奇分解,将其分解为一个势分量(共同目标可以表示)和一个调和分量(共同目标无法表示)。对于线性收益,在任意冲突图上和任意确定性平局打破协议下,两个分量均以闭式形式获得:以共同单位计,调和能量等于冲突数量,势能量则加上相邻冲突对的数量。因此,对于任意理性参数,智能体选择对数几率的最佳共同目标模型(在单边行动上均匀加权)的相对平方误差至少为$1/(d_{\max}+1)$,其中$d_{\max}$是单个智能体的最大冲突数量;对于八辆车交叉路口,该误差恰好为五分之一,无论优先级水平数量多少。对于严格改进动态不可见的被遗漏分量,在低理性对数线性学习(均匀修订)下,至首阶,它是平稳概率电流,其能量设定熵产生率。收益设计无法移除它:在$N$个智能体的完全冲突图上,在总序协议下且至少有三个优先级水平时,每个非常数基于秩的收益都会留下至少$1/N$的相对误差,且仅对于仿射收益达到等号。
英文摘要
In decentralised priority coordination, agents announce priority levels and a shared resource serves them in decreasing order, as at an unsignalised intersection; the levels form the decision layer of a hierarchical controller. Such interactions are routinely replaced by a potential game, i.e.\ by a common objective, for analysis and design. This paper determines what that surrogate misses, using the Hodge decomposition of the incentives into a potential component, which a common objective can represent, and a harmonic component, which it cannot. For the linear payoff, both components are obtained in closed form on every conflict graph and for every deterministic tie-breaking protocol: in common units, the harmonic energy is the number of conflicts and the potential energy adds the number of adjacent pairs of conflicts. Consequently, for every rationality parameter, the best common-objective model of the agents' choice log-odds, weighted uniformly over unilateral moves, has a relative squared error of at least $1/(d_{\max}+1)$, where $d_{\max}$ is the largest number of conflicts of one agent; for an eight-vehicle intersection it is exactly one fifth, for any number of priority levels. Invisible to strict-improvement dynamics, the missed component is, under low-rationality log-linear learning with uniform revision and to leading order, the stationary probability current, and its energy sets the entropy-production rate. Payoff design cannot remove it: on the complete conflict graph of $N$ agents, under a total-order protocol and with at least three priority levels, every nonconstant rank-based payoff leaves a relative error of at least $1/N$, with equality exactly for affine payoffs.