arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14327cs.LGstat.ML

非参数方差惩罚演员-评论家:风险敏感强化学习的统计推断

Nonparametric Variance-Penalized Actor-Critic: Statistical Inference for Risk-Sensitive Reinforcement Learning

Saunak Kumar Panda, Tong Li, Yisha Xiang, Ruiqi Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出非参数方差惩罚演员-评论家框架,用自举和随机缩放的在线估计器替代第二评论家,实现风险敏感强化学习的统计推断,并在HTS制造案例中显著降低变异性。

中文摘要 AI 辅助

方差惩罚是一种 principled 的风险敏感强化学习(RL)方法,它明确地在期望回报与策略稳定性之间进行权衡。现有方法需要专门的第二个评论家来在线估计回报方差,这增加了架构复杂性,并在学习过程中加剧了估计误差。我们提出了一种非参数方差惩罚演员-评论家(VPAC)框架,用基于自举和随机缩放的统计上可靠的在线估计器替代方差评论家,这些技术源自随机逼近的统计推断文献。这些估计器不需要辅助网络,保持单评论家架构,并且产生的方差惩罚在构造上是有界的,从而能够进行清晰的收敛性分析。我们通过常微分方程(ODE)方法,为方差惩罚Q学习算法和双时间尺度演员-评论家变体建立了几乎必然收敛性,仅要求方差估计保持有界而非一致。在实验上,我们在离散和连续随机环境中进行了评估,表明所提出的方法在方差减少方面达到或超过了现有双评论家VPAC基线,同时消除了第二个评论家的开销。我们进一步在高超导(HTS)制造案例研究中进行了验证,其中VPAC-RS(随机缩放)实现了稳态临界电流变异性降低74%,回合回报标准差降低63%,直接转化为改善的产量一致性。我们的结果确立了非参数统计推断作为风险敏感RL中辅助评论家的实用且理论上合理的替代方案。

英文摘要

Variance penalization is a principled approach to risk-sensitive reinforcement learning (RL) that explicitly trades expected return for policy stability. Existing methods require a dedicated second critic to estimate return variance online, adding architectural complexity and compounding estimation error during learning. We propose a nonparametric variance-penalized actor-critic (VPAC) framework that replaces the variance critic with statistically grounded online estimators based on bootstrapping and random scaling, techniques drawn from the statistical inference literature for stochastic approximation. These estimators require no auxiliary network, maintain a single-critic architecture, and produce variance penalties that are bounded by construction, enabling clean convergence analysis. We establish almost-sure convergence for both a variance-penalized Q-learning algorithm and a two-timescale actor-critic variant via the ordinary differential equation (ODE) method, requiring only that variance estimates remain bounded rather than consistent. Empirically, we evaluate across discrete and continuous stochastic environments, demonstrating that the proposed methods match or exceed the variance reduction achieved by the existing dual-critic VPAC baseline while eliminating the overhead of a second critic. We further validate on a high-temperature superconductor (HTS) manufacturing case study, where VPAC-RS (Random Scaling) achieves a 74% reduction in steady-state critical current variability and a 63% reduction in episode return standard deviation, translating directly to improved yield consistency. Our results establish nonparametric statistical inference as a practical and theoretically sound alternative to auxiliary critics for risk-sensitive RL.

发表机构

  • University of Houston(休斯顿大学)
  • Texas Tech University(德克萨斯理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑