arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2001.04032stat.MLcs.LG

POPCORN: Partially Observed Prediction COnstrained ReiNforcement Learning

  • Harvard SEAS(哈佛大学哈佛约翰·A·保尔森工程与应用科学学院)
  • Tufts University(塔夫茨大学)

机构由 AI 辅助整理,请以论文原文为准。

Joseph Futoma, Michael C. Hughes, Finale Doshi-Velez

更新

英文摘要:

Many medical decision-making tasks can be framed as partially observed Markov decision processes (POMDPs). However, prevailing two-stage approaches that first learn a POMDP and then solve it often fail because the model that best fits the data may not be well suited for planning. We introduce a new optimization objective that (a) produces both high-performing policies and high-quality generative models, even when some observations are irrelevant for planning, and (b) does so in batch off-policy settings that are typical in healthcare, when only retrospective data is available. We demonstrate our approach on synthetic examples and a challenging medical decision-making problem.

补充信息

↑