arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2504.00461cs.LGcs.GT

针对自适应对手的有向无环图在线最短路径问题在Bandit反馈下的高效近似最优算法

Efficient Near-Optimal Algorithm for Online Shortest Paths in Directed Acyclic Graphs with Bandit Feedback Against Adaptive Adversaries

  • University of Washington(华盛顿大学)
  • MIT(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Arnab Maiti, Zhiyuan Fan, Kevin Jamieson, Lillian J. Ratliff, Gabriele Farina

更新

AI总结:

本文研究了在自适应对手和Bandit反馈下有向无环图的在线最短路径问题,提出首个计算高效的算法,通过新颖的损失估计器和基于质心的分解实现了近似极小极大最优的遗憾界,并改进了多种博弈与老虎机问题的遗憾保证。

AI中文摘要:

本文研究在对抗自适应对手(adaptive adversary)且仅有Bandit反馈的情况下,有向无环图(DAGs)中的在线最短路径问题。给定一个包含源节点$v_{\mathsf{s}}$和汇节点$v_{\mathsf{t}}$的DAG $G = (V, E)$,令$X \subseteq \{0,1\}^{|E|}$表示从$v_{\mathsf{s}}$到$v_{\mathsf{t}}$的所有路径集合。在每一轮$t$中,我们选择一条路径$\mathbf{x}_t \in X$,并接收到关于损失$\langle \mathbf{x}_t, \mathbf{y}_t \rangle \in [-1,1]$的Bandit反馈,其中$\mathbf{y}_t$是对手选择的损失向量。我们的目标是最小化在$T$轮中相对于事后最佳路径的遗憾(regret)。我们提出了首个计算高效的算法,以高概率针对任何自适应对手实现了$\tilde O(\sqrt{|E|T\log |X|})$的近似极小极大最优遗憾界,其中$\tilde O(\cdot)$隐藏了关于边数$|E|$的对数因子。该算法以非平凡的方式利用了一种新颖的损失估计器和基于质心的分解(centroid-based decomposition)来达到此遗憾界。作为应用,我们证明了该DAG算法为$m$-sets、扩展型博弈、Colonel Blotto博弈、有向图中的最短游走、超立方体以及多任务多臂老虎机提供了目前最高效的算法,在所有这些场景中均实现了改进的高概率遗憾保证。

英文摘要:

In this paper, we study the online shortest path problem in directed acyclic graphs (DAGs) under bandit feedback against an adaptive adversary. Given a DAG $G = (V, E)$ with a source node $v_{\mathsf{s}}$ and a sink node $v_{\mathsf{t}}$, let $X \subseteq \{0,1\}^{|E|}$ denote the set of all paths from $v_{\mathsf{s}}$ to $v_{\mathsf{t}}$. At each round $t$, we select a path $\mathbf{x}_t \in X$ and receive bandit feedback on our loss $\langle \mathbf{x}_t, \mathbf{y}_t \rangle \in [-1,1]$, where $\mathbf{y}_t$ is an adversarially chosen loss vector. Our goal is to minimize regret with respect to the best path in hindsight over $T$ rounds. We propose the first computationally efficient algorithm to achieve a near-minimax optimal regret bound of $\tilde O(\sqrt{|E|T\log |X|})$ with high probability against any adaptive adversary, where $\tilde O(\cdot)$ hides logarithmic factors in the number of edges $|E|$. Our algorithm leverages a novel loss estimator and a centroid-based decomposition in a nontrivial manner to attain this regret bound. As an application, we show that our algorithm for DAGs provides state-of-the-art efficient algorithms for $m$-sets, extensive-form games, the Colonel Blotto game, shortest walks in directed graphs, hypercubes, and multi-task multi-armed bandits, achieving improved high-probability regret guarantees in all these settings.

补充信息

↑