Nonparametric Variance-Penalized Actor-Critic: Statistical Inference for Risk-Sensitive Reinforcement Learning
非参数方差惩罚演员-评论家:风险敏感强化学习的统计推断
机构 * University of Houston(休斯顿大学) ; Texas Tech University(德克萨斯理工大学)
AI总结 本文提出非参数方差惩罚演员-评论家框架,用自举和随机缩放的在线估计器替代第二评论家,实现风险敏感强化学习的统计推断,并在HTS制造案例中显著降低变异性。
Comments Submitted to IEEE Transactions on Neural Networks and Learning Systems. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible