适应去中心化异构老虎机中决策相关的非平稳性
Adapting to Decision-Relevant Non-Stationarity in Decentralized Heterogeneous Bandits
浏览论文内容
中文总结 AI 辅助
针对去中心化异构老虎机中的非平稳性,提出DRFC算法,通过平衡全局比较忽略局部变化,实现无局部变化依赖的遗憾界,并验证其有效性。
中文摘要 AI 辅助
去中心化老虎机系统通常包含异构智能体:即使网络的最佳动作保持不变,单个智能体的奖励也可能发生变化。当奖励在智能体间取平均时,这些局部变化可能相互抵消,因此局部变化的数量 $\Stloc$ 可能远大于最佳公共臂 $\Stdec$ 的变化数量。我们提出了决策相关的新鲜比较(DRFC),该方法使用来自所有智能体的新的、平衡的样本在网络层面比较臂,并且仅在新鲜的全局证据表明公共最佳臂已改变时才切换。我们证明了一个高概率动态遗憾界,其中不包含依赖于 $\Stloc$ 的适应项,并表明每个算法仍然必须为识别真正的决策切换并通过通信图传播它们而付出代价。在一种不同的时间平均基准下,一个随时有效的滑动窗口扩展可以处理逐渐漂移;在合成、半真实和 MovieLens-1M 重放上的实验表明,DRFC 忽略与决策无关的局部变化,而扩展避免了虚假切换。
英文摘要
Decentralized bandit systems often contain heterogeneous agents: rewards can change at individual agents even when the best action for the network stays the same. These local changes may cancel when rewards are averaged across agents, so the number of local changes $\Stloc$ can be much larger than the number of changes in the best common arm $\Stdec$. We introduce Decision-Relevant Fresh Comparison (DRFC), which uses new, balanced samples from all agents to compare arms at the network level and switches only when fresh global evidence indicates that the common best arm has changed. We prove a high-probability dynamic regret bound with no adaptation term depending on $\Stloc$, and show that every algorithm must still pay for identifying genuine decision switches and propagating them through the communication graph. Under a distinct time-average benchmark, an anytime-valid sliding-window extension handles gradual drift; experiments on synthetic, semi-real, and MovieLens-1M replays show that DRFC ignores decision-irrelevant local changes while the extension avoids false switches.