arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

随机微分方程引导的蒙特卡罗强化学习:一种用于噪声环境中稳健决策的随机最大值原理方法

SDE Guided Monte Carlo Reinforcement Learning: A Stochastic Maximum Principle Approach for Robust Decision Making in Noisy Environments

Juncai Wang

arXiv 2607.22541首次发表:更新:

AI 中文总结

研究探讨SMP定性最优条件能否指导稳健表格型强化学习算法,提出SDE-MC-AC框架,通过建立对应关系及自适应温度调度整合多种方法,经实验验证了相关假设,揭示导航模式,为随机最优控制与强化学习搭建桥梁。

AI 中文摘要

本文研究随机最大值原理(SMP)的定性最优条件能否作为设计稳健表格型强化学习算法的指导启发式方法。我们提出了SDE-MC-AC框架,建立了一组受SMP启发的对应关系。关键是,我们推导了一个受SMP启发的自适应温度调度,它根据局部值不确定性随机调整探索。我们在一个最小的固定迷宫测试平台上,在三种典型噪声模式下进行了一组五个有原则的实验,以测试从SMP得出的五个可证伪假设。结果为一些结论提供了证据,轨迹可视化揭示了与SDE模型一致的导航模式。该研究表明,将SMP启发式转换到离散MDP中能带来显著经验收益,在严格随机最优控制和实际稳健强化学习之间架起了桥梁。

英文摘要

This work investigates whether the qualitative optimality conditions of the SMP can serve as guiding heuristics for designing robust tabular RL algorithms. We propose the \textbf{SDE-MC-AC} framework, which establishes a set of SMP-inspired correspondences: the potential field gradient is mapped to potential-based reward shaping, the diffusion coefficient to a softmax temperature, and the value function gradient magnitude to the episode-averaged absolute temporal-difference (TD) error. Crucially, we derive an SMP-motivated adaptive temperature schedule that scales exploration stochastically in response to local value uncertainty. The resulting Monte Carlo Actor-Critic agent integrates potential-shaped rewards, entropy regularization, and adaptive temperature control within episodic updates. We conduct a set of five principled experiments in a minimal, fixed maze testbed under three canonical noise regimes (perceptual, dynamic goal, action) to test five falsifiable hypotheses derived from the SMP. Our results provide evidence for (i) an inverted-U optimal stochasticity curve, (ii) the robust failure prevention of the adaptive schedule under non-stationary noise, (iii) convergence acceleration by potential field guidance, (iv) cross-noise generalization, and (v) the indispensability of the combined SDE components through systematic ablation. Trajectory visualizations reveal a characteristic ``macroscopically deterministic, microscopically stochastic'' navigation pattern consistent with the SDE model. The study demonstrates that even a heuristic transposition of the SMP into a discrete MDP can yield significant empirical gains, thereby opening a bridge between rigorous stochastic optimal control and practical robust reinforcement learning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑