arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22012cs.LG

上下文博弈的跨域离策略评估与学习

Cross-Domain Off-Policy Evaluation and Learning for Contextual Bandits

Yuta Natsubori, Masataka Ushiku, Yuta Saito

首次发表
浏览论文内容

中文总结 AI 辅助

研究上下文博弈中离策略评估与学习问题,提出跨域OPE/L新设置,可利用目标域和其他域日志数据,开发新估计器和策略梯度方法,有效解决现有方法面临的挑战,增强离策略评估与学习能力。

中文摘要 AI 辅助

上下文博弈中的离策略评估与学习(OPE/L)在实际系统中迅速受到欢迎,因为仅使用历史记录数据就能安全地评估和学习新策略。然而,现有OPE/L方法无法处理诸如少样本数据、确定性日志记录策略和新动作等具有挑战性但普遍存在的场景。在许多应用中,我们需要在这些挑战下评估和学习新策略,现有方法因方差问题或日志数据中探索有限而无法有效评估和优化。为在这些未解决的挑战下实现OPE/L,我们提出跨域OPE/L的新问题设置,可访问目标域和其他域的日志数据。这种新颖的公式广泛适用,我们开发了新的估计器和策略梯度方法,利用目标和源数据集解决OPE/L,在实证评估中显著增强了OPE/L。

英文摘要

Off-Policy Evaluation and Learning (OPE/L) in contextual bandits is rapidly gaining popularity in real systems because new policies can be evaluated and learned securely using only historical logged data. However, existing methods in OPE/L cannot handle many challenging but prevalent scenarios such as few-shot data, deterministic logging policies, and new actions. In many applications, such as personalized medicine, content recommendations, education, and advertising, we need to evaluate and learn new policies in the presence of these challenges. Existing methods cannot evaluate and optimize effectively in these situations due to the notorious variance issue or limited exploration in the logged data. To enable OPE/L even under these unsolved challenges, we propose a new problem setup of Cross-Domain OPE/L, where we have access not only to the logged data from the target domain in which the new policy will be implemented but also to logged datasets collected from other domains. This novel formulation is widely applicable because we can often use historical data not only from the target hospital, country, device, or user segment but also from other hospitals, countries, devices, or segments. We develop a new estimator and policy gradient method to solve OPE/L by leveraging both target and source datasets, resulting in substantially enhanced OPE/L in the previously unsolved situations in our empirical evaluations.

发表机构

  • Hakuhodo DY Holdings, Inc.(博报堂DY控股公司)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑