AI 中文总结
针对多目标老虎机中多臂候选集的帕累托前沿近似问题,提出基于超体积贡献乐观估计的THV-UCB算法,定义近似超体积遗憾,给出无间隙和间隙相关的遗憾界,为小集合近似帕累托前沿提供理论支撑。
AI 中文摘要
我们研究一类随机多目标老虎机问题:在每一轮中,智能体选择包含$k$个臂的候选集,并在半老虎机反馈下观测这些臂的$d$维奖励向量。我们的目标并非识别单个最优臂,而是维护一个小规模动作集合,使其联合近似帕累托前沿。我们通过所选臂子集诱导的被支配超体积将该目标形式化,定义了相对于事后可获得的最优大小为$k$子集的$α$近似超体积遗憾,其中$α= 1 - 1/e$反映了单调次模函数贪心最大化的近似保证。为解决该问题,我们提出THV-UCB算法,这是一种基于臂的边际超体积贡献乐观估计进行贪心选臂的乐观算法。我们推导了适用于所有实例的无间隙遗憾界$\tilde{O}(d\sqrt{nkT})$,以及当臂充分分离时关于$T$呈多对数形式的间隙相关遗憾界$\tilde{O}(nk^{2.5}/Δ_{\min})$。我们的结果为各类多目标应用中使用小子集近似帕累托前沿提供了理论支撑。
英文摘要
We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback. We do not aim at identifying a single optimal arm; instead, we consider the problem of maintaining a small set of actions that jointly approximate the Pareto frontier. We formalize this objective through the dominated hypervolume induced by the selected subset of arms, and define an $α$-approximate hypervolume regret with respect to the best size-$k$ subset achievable in hindsight, where $α= 1 - 1/e$ reflects the approximation guarantee of greedy maximization for monotone submodular functions. To address this problem, we introduce \textit{THV-UCB}, an optimistic algorithm that selects arms greedily based on optimistic estimates of their marginal hypervolume contributions. We establish a gap-free regret bound $\tilde{O}(d\sqrt{nkT})$ that holds on every instance, together with a gap-dependent bound $\tilde{O}(nk^{2.5}/Δ_{\min})$ that becomes polylogarithmic in $T$ once the arms are sufficiently well separated. Our results provide theoretical support for using small subsets to approximate Pareto fronts in various multi-objective applications.
Comments21 pages, 7 figures, 7 tables