AI 中文总结
该研究提出 BC-ICL 方法,将预训练表格基础模型转化为在线决策随机策略,通过 Bootstrap 重采样与臂-上下文架构优化,在上下文多臂老虎机任务上取得优于基准的遗憾性能。
AI 中文摘要
上下文多臂老虎机为高效个性化提供了自然框架,但在稀疏、有偏的交互数据、不可靠的不确定性估计以及严重冷启动的情况下,实际部署仍然困难。我们研究是否可以将具备上下文学习(ICL)的预训练表格基础模型转化为用于在线决策的随机策略。我们提出 BC-ICL(使用 ICL 的 Bootstrap 条件动作选择),该方法在每一轮中对交互历史进行 Bootstrap 重采样,将冻结的预训练 ICL 模型以该重采样为条件,对所有动作进行评分,并选择采样得分最高的动作。我们还引入了一种臂-上下文条件架构,该架构可促进各动作间共享统计强度,有助于避免孤立臂多臂老虎机常见的 Bootstrap 失效模式。经验证,该策略在标准上下文多臂老虎机测试集上展现出优异的早期轮次遗憾和整体遗憾性能,在严格的在线协议下优于已有的基准方法。
英文摘要
Contextual bandits offer a natural framework for sample-efficient personalization, but practical deployment remains difficult under sparse, biased interaction data, unreliable uncertainty estimates, and severe cold starts. We study whether pre-trained tabular foundation models with in-context learning can be turned into randomized policies for online decision making. We propose BC-ICL (Bootstrap-conditioned action selection using ICL), which at each round draws a bootstrap resample of the interaction history, conditions a frozen pre-trained ICL model on that resample, scores all actions, and selects the action with the highest sampled score. We further introduce an arm-context conditioning architecture that promotes shared statistical strength across actions and helps avoid common bootstrap failure modes of isolated-arm bandits. Empirically, this policy delivers strong early-round regret and regret performance on standard contextual bandit suites, outperforming established baselines under a strict online protocol.