带探测的赌博机:最优遗憾与获胜反馈的极限
Bandits with Probing: Optimal Regret and the Limits of Winner Feedback
浏览论文内容
中文总结 AI 辅助
研究带探测的多臂赌博机问题,确定了在获胜反馈下极小极大遗憾的两个最优定律,分别对应独立随机奖励和固定序列,揭示了探测优势何时能覆盖学习成本。
中文摘要 AI 辅助
学习者每轮最多探测 $n$ 个臂中的 $k$ 个,获得它们在 $[0,1]$ 中奖励的最大值,并与最佳固定臂竞争。探测优势何时能为学习买单?我们确定了两个极小极大定律。在具有获胜反馈(最大值和获胜标签)的独立随机奖励下,或在给定块最大值之间的单个有符号对比的任意固定序列上,极小极大遗憾的阶为 $\Phi_{n,k}(T)=\min\{\frac{n-k}{n}T,\frac{n-k}{k}\}$,其中 $2\le k<n$。在获胜反馈下,任意联合独立同分布奖励和固定序列的极小极大遗憾阶均为 $R_{n,k}(T)=\frac{n-k}{n}\min\{T,\frac{n+T}{k},\sqrt{\frac{nT}{k}}\}$。两个定律都具有通用常数和任意时间上界。第一个定律将遗憾简化为纯覆盖成本:同轮对比吸收了稳定性成本,而独立性允许精确重采样,其收益为样本推进提供资金。第二个定律增加了一个学习成本,该成本在时间范围 $n$ 时变得与覆盖成本相当;超过 $nk$ 时,数值最大值比仅使用标签有所改进。下界允许任意自适应动作大小。
英文摘要
A learner probes at most $k$ of $n$ arms each round, receives the maximum of their rewards in $[0,1]$, and competes with the best fixed arm. When does the probing advantage pay for learning? We determine two minimax laws. Under independent stochastic rewards with winner feedback (the maximum and a winning label), or on arbitrary fixed sequences given a single signed contrast between block maxima, the minimax regret has order $Φ_{n,k}(T)=\min\{\frac{n-k}{n}T,\frac{n-k}{k}\}$, $2\le k<n$. Under winner feedback, both arbitrary joint i.i.d. rewards and fixed sequences have minimax regret of order $R_{n,k}(T)=\frac{n-k}{n}\min\{T,\frac{n+T}{k},\sqrt{\frac{nT}{k}}\}$. Both laws have universal constants and anytime upper bounds. The first reduces regret to a pure coverage cost: same-round contrasts absorb the stability cost, and independence permits exact resampling whose gains fund sample advancement. The second adds a learning cost that becomes comparable to coverage at horizon $n$; beyond $nk$, numerical maxima improve over labels alone. The lower bound allows every adaptive action size.
发表机构
- Zhejiang University of Technology(浙江工业大学)
机构由 AI 辅助整理,请以论文原文为准。