arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无参数重尾多臂老虎机

Parameter-Free Heavy-Tailed Bandits

Gianmarco Genalti, Alberto Maria Metelli

arXiv 2607.29460首次发表:更新:

发表机构

Politecnico di Milano(米兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究解决COLT 2025开放问题,提出无参数调度探索算法,实现未知重尾多臂老虎机的自适应,刻画了未知尾参数的统计代价。

AI 中文摘要

重尾分布自然出现在金融投资、在线广告和网络管理等序列决策问题中,其中罕见但极端的结果会主导性能。重尾多臂老虎机对这些场景下的在线决策进行建模,仅假设奖励$X$满足$\boldsymbol{E}[|X|^{1+\boldsymbol{\text{ε}}}]\boldsymbol{\text{≤}}\boldsymbol{u}$,其中尾指数$\boldsymbol{\text{ε}}\boldsymbol{\text{∈}}\boldsymbol{(0,1]}$,矩界$\boldsymbol{u}\boldsymbol{<}\boldsymbol{+}\boldsymbol{\text{∞}}$。然而,大多数现有遗憾最小化算法要求这些参数是已知的,这一假设在实践中极具局限性:$\boldsymbol{\text{ε}}$和$\boldsymbol{u}$分别控制罕见事件的频率和幅度,因此正是从有限观测中最难可靠推断的量。受Genalti和Metelli在COLT 2025上提出的开放问题的启发,我们解决了重尾多臂老虎机的无假设自适应问题,并刻画了未知尾参数带来的遗憾代价。我们首先研究固定尾指数$\boldsymbol{\text{ε}}$下对矩界$\boldsymbol{u}$的自适应,证明每一个未知$\boldsymbol{u}$或其任何上界的算法,都必须在其依赖分布和不依赖分布的遗憾保证之间遵循严格的权衡。随后,我们引入一种无需$\boldsymbol{u}$知识的 scheduled-exploration(调度探索)算法,该算法在对数因子内匹配所得的自适应前沿。最后,我们证明通过将探索调度校准到端点$\boldsymbol{\text{ε}}\boldsymbol{=}\boldsymbol{1}$,该算法可在未知$\boldsymbol{\text{ε}}$的情况下实例化,它对每一个固定$\boldsymbol{\text{ε}}\boldsymbol{>}\boldsymbol{0}$都能实现次线性遗憾,而没有任何算法能在所有$\boldsymbol{\text{ε}}\boldsymbol{\text{∈}}\boldsymbol{(0,1]}$上统一保证次线性遗憾。综上,我们的结果在无额外分布假设的情况下解决了COLT开放问题,并对适应未知重尾的统计代价给出了严格刻画。

英文摘要

Heavy-tailed distributions arise naturally in sequential decision-making problems such as financial investment, online advertising, and network management, where rare but extreme outcomes can dominate performance. Heavy-tailed bandits model online decision-making in these settings by assuming only that rewards $X$ satisfy $\mathbb{E}[|X|^{1+ε}]\leq u$, for some tail exponent $ε\in(0,1]$ and moment bound $u<+\infty$. However, most existing regret minimization algorithms require these parameters to be known. This assumption is particularly restrictive in practice: $ε$ and $u$ govern the frequency and magnitude of rare events and are therefore precisely the quantities that are hardest to infer reliably from limited observations. Motivated by an open problem posed by Genalti and Metelli at COLT 2025, we resolve the assumption-free adaptation problem for heavy-tailed bandits and characterize the price in the regret of not knowing the tail parameters. We first study adaptation to the moment bound $u$ for a fixed tail exponent $ε$. We prove that every algorithm unaware of $u$, or of any upper bound on it, must obey a sharp trade-off between its distribution-dependent and distribution-free regret guarantees. We then introduce a scheduled-exploration algorithm that requires no knowledge of $u$ and matches the resulting adaptation frontier up to logarithmic factors. Finally, we show that the same algorithm can be instanced without knowing $ε$ by calibrating its exploration schedule to the endpoint $ε=1$. It achieves sublinear regret for every fixed $ε>0$, while no algorithm can guarantee sublinear regret uniformly over all $ε\in(0,1]$. Altogether, our results resolve the COLT open problem without additional distributional assumptions and provide a sharp characterization of the statistical cost of adapting to unknown heavy tails.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑