发表机构
The Hong Kong University of Science and Technology; University of Pennsylvania(香港科技大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对带背包约束的稀疏线性上下文赌博机,提出抗腐败的ROPD和Shared-Grid算法,在奖励与消费腐败下实现亚线性遗憾并保持预算安全。
AI 中文摘要
我们研究了在联合奖励和消费腐败下的带背包约束的稀疏线性上下文赌博机。消费腐败带来的挑战超出了腐败奖励本身:它不仅影响统计估计,还影响记录的预算、资源价格以及控制未来分配的停止决策。我们开发了鲁棒乐观原始-对偶(ROPD)算法,这是一种估计器模块化框架,将腐败感知的置信宽度与在线资源价格和预算安全规则相结合。通过具体的稀疏实现,在强制探索和总体设计覆盖下,ROPD相对于干净的总体线性规划基准实现了$\tilde O(T^{2/3}+\Gamma T^{1/3})$的遗憾;在策略内实现设计覆盖下,实现了$\tilde O(\sqrt T+\Gamma)$的遗憾,其中$\Gamma$是提供的有效腐败界,并满足所述的比例预算缩放和固定模型/设计参数。当腐败水平未知时,共享网格(Shared-Grid)算法围绕拟合到单一实现历史的公共点估计调整置信半径,产生显式的初始化和主比较成本;其更精细的策略内保证额外需要推荐覆盖。两种方法在每次实现中都保持观察到的预算,并通过累积消费腐败来限制清洁资源违规。这些结果将抗腐败的稀疏估计与高维在线分配中的资源核算、定价和停止联系起来。
英文摘要
We study sparse linear contextual bandits with knapsack constraints under joint reward and consumption corruption. Consumption corruption creates a challenge beyond corrupted rewards: it affects not only statistical estimates, but also the recorded budget, resource prices, and stopping decisions that govern future allocation. We develop Robust Optimistic Primal--Dual (ROPD), an estimator-modular framework that combines corruption-aware confidence widths with online resource prices and a budget-safety rule. With concrete sparse implementation, ROPD achieves regret against a clean population-LP benchmark of $\widetilde O(T^{2/3}+ΓT^{1/3})$ under forced exploration and population-design coverage, and $\widetilde O(\sqrt T+Γ)$ under on-policy realized-design coverage, for a supplied valid corruption bound $Γ$ under the stated proportional-budget scaling and fixed model/design parameters. When the corruption level is unknown, Shared-Grid adapts confidence radii around common point estimates fitted to a single realized history, incurring explicit initialization and master-comparison costs; its sharper on-policy guarantee additionally requires recommendation coverage. Both methods preserve observed budgets on every realization and bound clean resource violation by cumulative consumption corruption. These results connect corruption-robust sparse estimation with resource accounting, pricing, and stopping in high-dimensional online allocation.