发表机构
The University of Osaka(大阪大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证明连续时间风险敏感控制等价于Rényi散度最小化,通过Girsanov定理将风险敏感度映射为散度阶数,统一了任意风险态度的概率控制框架,为采样控制算法奠定基础。
AI 中文摘要
在本研究中,我们证明了一个连续时间风险敏感控制问题等价于在轨迹路径测度上的Rényi散度最小化问题。通过Kullback-Leibler(KL)散度最小化将随机最优控制重新表述为概率推断,避免了Hamilton-Jacobi-Bellman方程的计算难解性。然而,标准的KL控制本质上是风险中性的,而最近的极小极大扩展仍局限于风险规避设置。我们的等价性结果通过为连续时间非线性系统中的任意风险态度提供统一的概率框架,解决了这一局限性。基于Girsanov定理,我们明确地将风险敏感度映射到Rényi散度阶数,推导出由风险偏好加权的依赖于噪声的控制惩罚项。该公式无缝地调节尾部加权行为,在风险规避策略的零强制与风险寻求策略的质量覆盖之间进行插值。这些发现弥合了随机控制与信息论推断之间的鸿沟,为基于采样的控制算法提供了基础。
英文摘要
In this study, we show that a continuous-time risk-sensitive control problem is equivalent to a Rényi divergence minimization problem over trajectory path measures. Reformulating stochastic optimal control as probabilistic inference via Kullback-Leibler (KL) divergence minimization avoids the computational intractability of the Hamilton-Jacobi-Bellman equation. However, standard KL control is inherently risk-neutral, and recent minimax extensions remain restricted to risk-averse settings. Our equivalence result resolves this limitation by offering a unified probabilistic framework for arbitrary risk attitudes in continuous-time nonlinear systems. Based on Girsanov theorem, we explicitly map the risk sensitivity to the Rényi divergence order, deriving a noise-dependent control penalty scaled by risk preference. This formulation seamlessly modulates tail-weighting behaviors, interpolating between zero-forcing for risk-averse policies and mass-covering for risk-seeking policies. These findings bridge stochastic control and information-theoretic inference, providing a foundation for sampling-based control algorithms.
Comments6 pages, 1 figure