arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规划即推理的对偶平坦几何

The Dually Flat Geometry of Planning as Inference

Nikola Milosevic, Asaki Kataoka, Nicolas Hinrichs, Kenji Doya, Nico Scherf

arXiv 2609.04005首次发表:更新:

发表机构

Max Planck Institute for Human Cognitive and Brain Sciences; Okinawa Institute of Science and Technology(马克斯·普朗克人类认知与脑科学研究所; 冲绳科学技术大学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出强化学习中占用度的新表征,定义访问度并构建其对偶平坦统计流形,推广规划即推理至非线性泛函,推导其对强化学习和理论神经科学的意义。

AI 中文摘要

我们提出了强化学习中占用度的另一种表征,该表征通过重置规划过程将规划准则嵌入动力学而得到。其平稳测度(我们称之为访问度)是决策信息几何最自然的表达对象。可实现的访问度构成一个对偶平坦统计流形,其两个仿射图分别为访问概率和对数策略,二者在条件熵下对偶。该结构使规划即推理能从线性奖励推广到访问度的非线性泛函,每次迭代通过一次自然梯度步求解,并将时间差分误差解释为边际效用估计。我们推导了该几何及其对强化学习和理论神经科学的意义。

英文摘要

We present an alternative characterization of the occupancy measure of reinforcement learning, obtained by embedding the planning criterion into the dynamics through a resetting planning process. Its stationary measure, which we term visitation measure, is the object on which the information geometry of decision making is most naturally expressed. The achievable visitation measures form a dually flat statistical manifold whose two affine charts are the visitation probabilities and the log-policies, dual under the conditional entropy. This structure makes planning-as-inference generalize from linear rewards to nonlinear functionals of the visitation, each iterate solved by one natural-gradient step, and gives the temporal-difference error the interpretation of a marginal-utility estimate. We develop the geometry and its consequences for reinforcement learning and theoretical neuroscience.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑