arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从最优动作到世界模型:折扣马尔可夫决策过程中转移核的可识别性

From Optimal Actions to World Models: Identifiability of Transition Kernels in Discounted MDPs

Neal Batra

arXiv 2608.07301首次发表:更新:

AI 中文总结

该研究探讨折扣马尔可夫决策过程中仅从最优动作可恢复的转移概率信息,明确了不同形式奖励对转移核可识别性的影响,证明了转移核的可识别条件及相关维度特性。

AI 中文摘要

我们研究仅从最优动作中可以恢复马尔可夫决策过程(MDP)的转移概率的哪些信息,这与Letcher等人考虑的逆问题密切相关,Letcher等人研究何时能从数值Q值中恢复动态。在本研究中,并未观测到数值本身,仅在给定类别的每个奖励下的最优动作是已知的。对于状态-动作奖励r(s,a),知道每个奖励下的最优动作还能告知我们,当每个动作都遵循相同的固定策略时,一个动作比另一个动作好多少,但这仍不足以唯一确定转移概率。我们证明,当存在一个满足L1=1的可逆矩阵L时,两个核会在每个奖励下给出相同的最优动作,公式为Q_{s,a}=(P_{s,a}+(1/γ)e_s^T(L-I))L^{-1}。在具有严格正项的核附近,存在一个n(n-1)维的不同核族具有此性质。若仅考虑在每个状态都有唯一最优动作的奖励,结果不变。随后我们将此与形式为r(s)和r(s,a,s')的奖励进行比较:依赖于下一个状态的奖励通常可以恢复转移核本身,每个至少有两个动作的状态的行都可被确定,我们还明确描述了当状态只有一个动作时,该行何时仍保持隐藏;状态奖励揭示的信息更少,两个核给出相同最优动作的情况恰好是每个确定性策略对同一组奖励都是最优的。这些结果表明,奖励的形式如何影响仅从最优动作中可学习到的动态信息。

英文摘要

We study what can be recovered about the transition probabilities of a Markov decision process from optimal actions alone. This is closely related to the inverse problem considered by Letcher et al., who ask when the dynamics can be recovered from numerical \(Q\)-values. Here the numerical values themselves are not observed; only the optimal actions are known, for every reward in a given class. For state-action rewards \(r(s,a)\), knowing the optimal actions for every reward also tells us how much better one action is than another when each is followed by the same fixed policy. This is still not enough to determine the transition probabilities uniquely. We prove that two kernels give the same optimal actions for every reward exactly when \[ Q_{s,a} = \Bigl(P_{s,a}+\tfrac1γe_s^{\mathsf T}(L-I)\Bigr)L^{-1} \] for one invertible matrix \(L\) satisfying \(L\mathbf 1=\mathbf 1\). Near a kernel with strictly positive entries, there is an \(n(n-1)\)-dimensional family of different kernels with this property. The result is unchanged if we consider only rewards having a unique optimal action at every state. We then compare this with rewards of the forms \(r(s)\) and \(r(s,a,s')\). Rewards that depend on the next state can usually recover the transition kernel itself: every row at a state with at least two actions is determined, and we describe exactly when a row at a state with one action can remain hidden. State rewards reveal less: two kernels give the same optimal actions exactly when every deterministic policy is optimal for the same set of rewards. The results show how the form of the reward affects what can be learned about the dynamics from optimal actions alone.

Comments15 pages, 1 figure

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑