arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向大动作空间的在线策略与离线策略学习

On-Policy and Off-Policy Learning for Large Action Spaces

Imad Aouali

arXiv 2607.28408首次发表:更新:

AI 中文总结

本论文针对大动作空间交互式系统的策略学习挑战,提出meTS、dTS、sDM等结构化贝叶斯方法,解决在线与离线策略学习的关键问题并提供理论保证。

AI 中文摘要

本论文研究交互式系统中的策略学习,其中智能体观察上下文、从极大集合中选择动作并接收部分反馈,主要框架为上下文多臂老虎机,包含在线策略学习与离线策略学习两大范式:在线策略学习指智能体与环境顺序交互并最小化遗憾,离线策略学习指其从记录策略收集的日志数据中学习。在大动作空间中,两种设置均面临低效探索、稀疏数据覆盖、高方差重要性权重、外推偏差及优化空间复杂等重大挑战。第一部分开发在线策略学习的结构化贝叶斯方法,引入混合效应汤普森采样扩展方法meTS,以及利用扩散启发先验建模动作间依赖关系的dTS,这些方法跨动作共享信息并根据有效动作数量提供遗憾保证。第二部分解决离线策略学习问题,提出基于隐变量的结构化直接方法sDM,证明大动作空间中优化误差可主导估计误差,并引入凹的、可高效优化的策略加权对数似然目标。最后,开发基于指数平滑和PAC-贝叶斯界的可微悲观方法,以控制正则化重要性采样估计器的偏差-方差权衡。

英文摘要

This thesis studies policy learning in interactive systems where an agent observes a context, selects an action from a very large set, and receives partial feedback. The main framework is contextual bandits, with two paradigms: on-policy learning, where the agent interacts sequentially with the environment and minimizes regret, and off-policy learning, where it learns from logged data collected by a logging policy. In large action spaces, both settings face major challenges: inefficient exploration, sparse data coverage, high-variance importance weights, extrapolation bias, and difficult optimization landscapes. The first part develops structured Bayesian methods for on-policy learning. We introduce meTS, a mixed-effect extension of Thompson sampling, and dTS, which leverages diffusion-inspired priors to model dependencies between actions. These methods share information across actions and yield regret guarantees depending on an effective number of actions. The second part addresses off-policy learning. We propose sDM, a structured direct method based on latent variables, show that optimization error can dominate estimation error in large action spaces, and introduce concave, efficiently optimizable policy-weighted log-likelihood objectives. Finally, we develop differentiable pessimistic methods based on exponential smoothing and PAC-Bayesian bounds to control the bias-variance trade-off of regularized importance-sampling estimators.

CommentsPhD Thesis, 241 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑