arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07998cs.AImath.OC

小批量风险厌恶深度Q学习:机器人导航案例研究

Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study

Aayush Patel, Andrzej Ruszczyński

首次发表
浏览论文内容

中文总结 AI 辅助

针对风险度量与强化学习结合难的问题,提出小批量转移风险映射嵌入双深度Q网络的风险厌恶Q学习方法,应用于水下机器人导航,在保留环境实验中提升性能并降低风险。

中文摘要 AI 辅助

我们研究马尔可夫决策过程的控制问题,其中策略的质量通过动态、时间一致的马氏风险度量而非期望折扣成本来评估。将此类度量与强化学习结合的主要障碍在于,转移风险映射以非线性方式依赖于转移核,因此无法从单个观测转移中估计。我们通过采用小批量转移风险映射来消除这一障碍:该映射应用于$N$个独立下一状态样本的经验度量,并对结果取平均。所得映射再次具有一致性。然而,作为$N$个下一状态值函数的期望值,它允许无偏的单样本估计器。我们将此映射嵌入双深度Q网络,分析由此产生的两个估计偏差来源,并获得一种适用于远超表格方案可及状态空间的风险厌恶Q学习方法。该方法应用于水下机器人导航问题,其中车辆必须访问收集点,收集随机信息载荷,并将其交付到传输点,同时每一步都面临被摧毁的风险。分层分解将路径执行委托给精确图搜索,并将学习限制在高层的“收集或传输”决策。一个低维特征映射,在问题的对称性下不变,取代了原始的状态-配置编码。在$300$个保留环境上的实验中,所得策略泛化到训练中从未见过的实例规模,并且当模拟器被错误指定时,仅$N=2$就降低了结果分布的上半偏差,同时改善了其均值——这是一致性风险度量与分布鲁棒性之间对偶性的经验对应。

英文摘要

We study the control of Markov decision processes in which the quality of a policy is evaluated by a dynamic, time-consistent Markov risk measure rather than by an expected discounted cost. The main obstacle to combining such measures with reinforcement learning is that a transition risk mapping depends on the transition kernel in a nonlinear way, and therefore cannot be estimated from a single observed transition. We remove this obstacle by employing mini-batch transition risk mappings: the mapping is applied to the empirical measure of $N$ independent next-state samples, and the result is averaged. The resulting mapping is again coherent. However, as an expected value of a function of $N$ next-state values, it admits an unbiased one-sample estimator. We embed this mapping into a double deep Q-network, analyze the two sources of estimation bias that arise, and obtain a risk-averse Q-learning method applicable to state spaces far beyond the reach of tabular schemes. The method is applied to an underwater robot navigation problem, in which a vehicle must visit collection points, gather stochastic information payloads, and deliver them at transmission points, while exposed at each step to the risk of destruction. A hierarchical decomposition delegates path execution to an exact graph search and confines learning to the high-level ``collect or transmit'' decision. A low-dimensional feature map, invariant under the symmetries of the problem, replaces the raw state--configuration encoding. In experiments on $300$ held-out environments, the resulting policies transfer to instance sizes never seen in training, and already $N=2$ reduces the upper semideviation of the outcome distribution while simultaneously improving its mean whenever the simulator is misspecified---an empirical counterpart of the duality between coherent risk measures and distributional robustness.

发表机构

  • Rutgers University(罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑