arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

因果博弈中的信息导向采样

Information-Directed Sampling for Causal Bandits

Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi

arXiv 2607.15577首次发表:更新:

AI 中文总结

研究含不可操纵变量的情境因果博弈,假设因果图已知且无潜在混杂,采用贝叶斯公式。开发汤普森采样和信息导向采样的因果变体,建立相关遗憾界和置信界,实验证明所提方法能有效利用信息,优于因果和非因果基线。

AI 中文摘要

因果博弈利用变量间的结构关系在干预间共享信息,加速高回报决策识别。然而在许多应用中,一些变量虽影响奖励但无法直接操纵。本文研究含不可操纵变量的情境因果博弈,行动选择前观察情境变量,每次干预后观察额外变量。假设因果图已知且无潜在混杂,采用贝叶斯公式,观测分布的条件概率表为未知参数。据此开发了汤普森采样和信息导向采样的因果变体。为汤普森采样建立了依赖熵的次线性贝叶斯遗憾界,为信息导向采样导出了依赖熵的遗憾界,还给出算法中蒙特卡罗估计的高概率置信界。实验表明所提方法通过有效利用干预间共享信息优于因果和非因果基线。

英文摘要

Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applications, however, some variables cannot be directly manipulated, even though they influence the reward and provide useful information about the underlying causal system. We study contextual causal bandits with non-manipulable variables, where context variables are observed before action selection and additional variables are observed after each intervention. Assuming a known causal graph without latent confounding, we adopt a Bayesian formulation in which the conditional probability tables of the observational distribution constitute the unknown parameter. This representation allows observations collected under one intervention to update reward estimates for other interventions through their shared causal mechanisms. We develop causal variants of Thompson Sampling and Information-Directed Sampling (IDS) for this setting. For Thompson Sampling, we establish an entropy-dependent sublinear Bayesian regret bound. For IDS, we derive an entropy-dependent regret bound that explicitly quantifies the additional error introduced by Monte Carlo approximation of the expected regret and information gain; when these quantities are available exactly, the bound recovers the standard sublinear IDS rate. The dependence of these guarantees on the action-set size is worst-case: our model contains the standard multi-armed bandit as a special case. We further provide high-probability confidence bounds for the Monte Carlo estimates. Experiments on synthetic causal bandit tasks show that the proposed methods outperform causal and non-causal baselines by effectively exploiting information shared across interventions.

CommentsAccepted for publication in TMLR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑