arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使玻尔兹曼理性合理化:熵正则化策略的公理表征

Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies

Silviu Pitis

arXiv 2607.17316首次发表:更新:

AI 中文总结

研究强化学习中softmax策略,通过区分两种随机性化解其与MDP奖励结构独立性公理的矛盾,施加IIA和单调性确定玻尔兹曼策略等,得出RL相关结果,综合多领域线索对IIA用于智能体设计进行规范评估。

AI 中文摘要

在强化学习(RL)中,softmax策略$\pi(a \mid s) \propto \exp(\beta Q(s,a))$是随机选择的默认模型。RL文献中基于鲁棒性、探索和优化给出了各种理由,但都没有从第一原理唯一推导出softmax形式。这使得一个基本矛盾未解决:软贝尔曼方程中的熵奖励违反了支撑马尔可夫决策过程(MDP)奖励结构的独立性公理。我们通过区分机会和选择这两种随机性来化解这一矛盾。通过将冯·诺依曼 - 摩根斯坦(VNM)独立性限制在基本前景的环境彩票上,我们表明在选择节点对策略和价值函数施加无关选项独立性(IIA)和单调性唯一地确定了玻尔兹曼策略、熵正则化表示和软贝尔曼方程。软贝尔曼方程和硬贝尔曼方程之间的选择因此归结为一个设计决策:智能体是否重视自己的选择能力。我们得出了RL特有的结果,包括回报单调性和广义折扣下的收敛性,并综合了经济学和信息理论中得出相同结构的独立线索,对IIA何时适用于智能体设计进行了规范评估。

英文摘要

The softmax policy $π(a \mid s) \propto \exp(βQ(s,a))$ is the default model of stochastic choice in reinforcement learning (RL). Various justifications based on robustness, exploration, and optimization have been offered in the RL literature, but none uniquely derives the softmax form from first principles. This leaves a basic tension unresolved: the entropy bonus in the soft Bellman equation violates the Independence axiom that underwrites the Markov decision process (MDP) reward structure. We dissolve this tension by distinguishing two kinds of randomness: chance and choice. By restricting von Neumann-Morgenstern (VNM) Independence to environmental lotteries over base prospects, we show that imposing independence of irrelevant alternatives (IIA) and monotonicity on the policy and value functions at choice nodes uniquely determines the Boltzmann policy, the entropy-regularized representation, and the soft Bellman equation. The choice between the soft and hard Bellman equations thus reduces to a design decision: whether the agent values its own ability to choose. We develop RL-specific consequences, including return monotonicity and convergence under generalized discounting, and synthesize the independent lines from economics and information theory that arrive at the same structure, offering a normative assessment of when IIA is appropriate for agent design.

CommentsRLC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑