arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10981cs.LG

非单调凸岭赌博机的汤普森采样:多项式遗憾不需要单调性

Thompson Sampling for Non-Monotone Convex Ridge Bandits: Monotonicity Is Not Needed for Polynomial Regret

Xuan Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文证明汤普森采样在非单调凸岭赌博机中仍能达到多项式贝叶斯遗憾,通过新的基数界替代失效的椭球二分法,回答了单调性是否必要的问题。

中文摘要 AI 辅助

Bakhtiari、Lattimore和Szepesvári(COLT 2025)证明了对于具有凸单调岭损失$f(x)=\ell(\ip{x}{\theta})$的赌博机凸优化,汤普森采样(TS)具有贝叶斯遗憾$\tilde O(d^{5/2}\sqrt n)$,并询问链接的单调性是否是必要的。我们给出了一个定性的否定答案。对于取值于$[0,1]$的$1$-Lipschitz凸岭损失的任意先验,且链接为任意凸的、可能非单调的链接,以及任何固定的可测最小化器选择,精确后验TS具有贝叶斯遗憾$O\big((d+1)^4\sqrt{dn}\\,\log(e+nd\max\{1,\diam K\})\big)=\tilde O(d^{9/2}\sqrt n)$。单调情形的证明依赖于单次移除的John椭球二分法;我们通过一个明确的十二点配置表明该二分法对非单调链接失效,并用“无信息”配置的$O(d^2)$基数界替代它。该界使用布尔舍入论证:一个在最大范数下与秩为$r$的矩阵距离在$1/(4r)$以内的$0$-$1$矩阵的秩至多为$2r-1$。我们构造了$d(d+1)$个无信息损失,表明在大直径与间隙之比区域中该基数界在常数意义下是紧的,并给出了一个自包含的信息比到遗憾的转换,该转换对固定的可测选择是一致的。单调情形中$d^{5/2}$的依赖是否能够保留仍然是一个开放问题。

英文摘要

Bakhtiari, Lattimore and Szepesvári (COLT 2025) proved that Thompson sampling (TS) has Bayesian regret $\tilde O(d^{5/2}\sqrt n)$ for bandit convex optimisation with convex \emph{monotone} ridge losses $f(x)=\ell(\ip{x}θ)$, and asked whether monotonicity of the link is necessary. We give a qualitative negative answer. For every prior on $[0,1]$-valued, $1$-Lipschitz convex ridge losses with an arbitrary convex, possibly non-monotone, link, and for any fixed measurable selection of minimisers, exact-posterior TS has Bayesian regret $O\big((d+1)^4\sqrt{dn}\,\log(e+nd\max\{1,\diam K\})\big)=\tilde O(d^{9/2}\sqrt n)$. The monotone proof relies on a single-removal John-ellipsoid dichotomy; we show by an explicit twelve-point configuration that this dichotomy fails for non-monotone links, and replace it by an $O(d^2)$ cardinality bound for ``uninformative'' configurations. The bound uses a Boolean rounding argument: a $0$-$1$ matrix within $1/(4r)$ in max-norm of a rank-$r$ matrix has rank at most $2r-1$. We construct $d(d+1)$ uninformative losses, showing that the cardinality bound is tight up to constants in the large-diameter-to-gap regime, and give a self-contained information-ratio-to-regret transfer that is uniform over fixed measurable selections. Whether the $d^{5/2}$ dependence of the monotone case can be retained remains open.

发表机构

  • University of New South Wales(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

↑