arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25241stat.MLcs.LG

从未见过的角度学习:具有隐藏动作的离线强化学习

Learning from the Unseen: Offline Reinforcement Learning with Hidden Actions

Zeyu Bian, Ying Zhou, Yifan Cui

首次发表
浏览论文内容

中文总结 AI 辅助

研究具有隐藏动作的离线强化学习问题,利用下一状态变量为未观测动作的代理,提出LURE估计器,该估计器多重稳健且渐近正态,通过模拟和败血症管理应用证明了其有效性。

中文摘要 AI 辅助

标准的离线强化学习算法通常假定数据集中的动作能被无误差观测到。但在许多实际应用中,真实动作无法观测,只有有噪声的代理可用,这会使现有强化学习方法得出有偏差且可能误导性的结论。我们研究了具有隐藏动作的无限期折扣马尔可夫决策过程中的离策略评估。通过利用下一状态变量作为未观测动作的自然代理,我们确定了策略值并提出了一种基于影响函数的估计器LURE。LURE具有多重稳健性,在几种正确指定的干扰成分组合下保持一致性,且渐近正态,能进行有效的统计推断。据我们所知,这是第一项解决具有隐藏动作的离线强化学习的工作。我们通过模拟以及使用MIMIC - III数据库的败血症管理应用展示了LURE的有效性。

英文摘要

Standard offline reinforcement learning (RL) algorithms typically assume that the actions in the dataset are observed without error. However, in many real-world applications, the true actions are unobserved and only noisy proxies are available, causing existing RL methods to yield biased and potentially misleading conclusions. We study off-policy evaluation in infinite-horizon discounted Markov decision processes with hidden actions. By leveraging the next-state variable as a natural proxy for the unobserved action, we establish identification of the policy value and propose an influence-function-based estimator called LURE (Learning from the Unseen: Robust Estimator). LURE is multiply robust, remaining consistent under several combinations of correctly specified nuisance components, and is asymptotically normal, enabling valid statistical inference. To our knowledge, this is the first work to address offline RL with hidden actions. We demonstrate LURE's effectiveness through simulations and a sepsis management application using the MIMIC-III database.

发表机构

  • Florida State University(佛罗里达州立大学)
  • University of Connecticut(康涅狄格大学)
  • Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

↑