当贪心采样进行探索:无 Eluder 维度依赖的 KL 正则化上下文赌博机
When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence
浏览论文内容
中文总结 AI 辅助
本文研究 KL 正则化上下文赌博机,证明贪心采样在奖励与偏好反馈下均可实现无 eluder 维度依赖的对数遗憾,并揭示其与上置信界探索间的权衡。
中文摘要 AI 辅助
我们研究了在奖励反馈和偏好反馈两种设置下的 KL 正则化上下文赌博机问题。我们证明贪心采样可以在不显式依赖 eluder 维度的前提下实现对数遗憾。对于奖励反馈,我们为一个简单的贪心算法建立了与 eluder 维度无关的遗憾界,该算法直接从由估计奖励诱导的 Gibbs 策略中采样。我们进一步将此结果推广到一般偏好模型和 Bradley-Terry 模型下的偏好反馈,同时改进了现有的依赖维度的保证。我们的分析揭示了贪心采样与上置信界风格探索之间的权衡:当 KL 正则化足够强时,贪心采样享有更强的保证;而当正则化减弱时,额外的探索变得更为可取。
英文摘要
We study KL-regularized contextual bandits under both reward and preference feedback. While existing regret guarantees typically depend on the eluder dimension, we show that simple greedy sampling can achieve polylogarithmic regret without explicit dependence on this complexity measure. For reward feedback, we analyze a greedy algorithm that samples directly from the Gibbs policy induced by the estimated reward. We extend the result to preference feedback under both general preference and Bradley--Terry models, while also sharpening existing dimension-dependent guarantees. Our analysis reveals a trade-off between greedy sampling and upper confidence bound-style exploration: greedy sampling enjoys stronger regret guarantees when KL regularization is sufficiently strong, whereas additional exploration yields sharper bounds as the regularization weakens.
发表机构
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- Oregon State University(俄勒冈州立大学)
机构由 AI 辅助整理,请以论文原文为准。