arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

替代指标不是回报:上下文赌博机中替代指标后的主要结果采集

The Surrogate Is Not the Reward: Post-Surrogate Primary-Outcome Acquisition in Contextual Bandits

Kyungbok Lee, Michael R. Kosorok

arXiv 2610.06610首次发表:更新:

发表机构

University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出审计替代指标赌博机(ASB),在上下文赌博机中基于决策相关性和残余不确定性分配主要结果采集预算,实现亚线性遗憾并优于单一因素方法。

AI 中文摘要

我们研究上下文赌博机,其中在采取行动之后、学习者决定是否采集定义行动价值和遗憾的主要结果之前,会观察到替代指标。采集主要结果的价值取决于决策相关性(当前结果对比较策略的重要程度)以及观察到替代指标后的残余不确定性。审计替代指标赌博机(ASB)在T轮中分配B次主要结果采集预算的同时学习上下文策略。ASB根据当前决策相关性设定采集前的水平,并在观察到替代指标后,利用残余不确定性的估计值重新分配该水平。对于K个动作上的N个策略的有限类别,ASB相对于类别中最优策略产生\widetilde O[\sqrt{KT\log N}\{1+\sqrt{T/B}\}]的遗憾。在一个双动作家族中,当替代指标不能揭示更优动作时,在决定是否采集之前观察到替代指标的学习者可以实现有界遗憾,而任何必须在看到替代指标之前决定的学习者在相同预算下会产生\Omega(T/B)的最坏情况遗憾。合成实验表明,两个采集因素都很重要:ASB比仅使用决策相关性或仅使用残余不确定性的变体具有更低的遗憾。在KuaiRec用户-视频交互基准上,相对于仅相关性采集的遗憾差距随着预算增长先扩大后缩小。

英文摘要

We study contextual bandits in which a surrogate is observed after the action but before the learner decides whether to acquire the primary outcome that defines action value and regret. The value of acquiring the primary outcome depends on both decision relevance (how much the current outcome matters for comparing policies) and the residual uncertainty after observing the surrogate. The Audited Surrogate Bandit (ASB) learns a contextual policy while allocating a budget of $B$ primary-outcome acquisitions over $T$ rounds. ASB sets a pre-surrogate acquisition level from current decision relevance and, after observing the surrogate, redistributes that level using an estimate of that residual uncertainty. For a finite class of $N$ policies over $K$ actions, ASB incurs $\widetilde O[\sqrt{KT\log N}\{1+\sqrt{T/B}\}]$ regret relative to the best policy in the class. In a two-action family where the surrogate does not reveal the better action, a learner that observes the surrogate before deciding whether to acquire can achieve bounded regret, whereas any learner that must decide before seeing the surrogate incurs $Ω(T/B)$ worst-case regret under the same budget. Synthetic experiments show that both acquisition factors matter: ASB has lower regret than variants using only decision relevance or only residual uncertainty. On a KuaiRec benchmark of user-video interactions, the regret gap relative to relevance-only acquisition widens and then narrows as the budget grows.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑