arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

广义贝尔曼递归与序贯决策中的三种对偶性

Generalised Bellman recurrence and three dualities in sequential decision-making

Fernando E. Rosas, David Hyland, Daniel Polani

arXiv 2607.18077首次发表:更新:

AI 中文总结

研究贝尔曼方程形式的来源,表明最优值函数递归性质源于三个条件,这些条件产生三种对偶性,框架揭示对偶性源于单一构造,统一了强化学习、控制和决策理论中相关方法。

AI 中文摘要

贝尔曼方程的形式是由何而来?我们表明,最优值函数的递归性质源于三个条件:动力学通过充分统计量分解、回报递归分解、不确定性的聚合与前两者兼容。当这三个条件在共同状态下都成立时,贝尔曼方程源于它们的相互一致性;当其中一个不成立时,通常可通过扩充状态或变形回报或动力学来恢复可处理性。相同条件还产生了三种对偶性:概率与回报之间的对偶性、回报与聚合之间的对偶性、聚合与概率之间的对偶性。我们的框架揭示这些对偶性源于单一构造,统一了强化学习、控制和决策理论中分别发展的方法。

英文摘要

What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recursively, and that the aggregation of uncertainty is compatible with both. When all three conditions hold on a common state, the Bellman equation arises from their mutual consistency; when one fails, tractability can often be recovered by augmenting the state or by deforming return or dynamics. The same conditions are shown to give rise to three dualities: one between probability and return, one between return and aggregation, and one between aggregation and probability. Our framework reveals these dualities as arising from a single construction, unifying methods developed separately across reinforcement learning, control, and decision theory.

Comments27 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑