AI 中文总结
研究有限时域动态规划中依赖历史偏好的马尔可夫决策过程,通过行为公理等得出规范PA状态及递归表示,在一定条件下诱导最优策略,附加条件可重新参数化偏好记忆,给出示例分类说明框架范围。
AI 中文摘要
在具有依赖历史偏好的有限时域动态规划中,相关状态可能是整个已实现历史,即便物理状态是马尔可夫的。本文为这类马尔可夫决策过程发展了一种行为状态约简理论。在行为公理和确定性等价丰富条件下,全历史问题有由时间和风险聚合器组成的递归表示。通过对具有相同当前物理马尔可夫状态、在每个共同延续计划下无差异且在每个共同一步扩展后仍等价的历史进行商运算,得出一个规范的偏好增强(PA)状态。该规范PA状态在基础偏好的可达递归分解中是最小的。在马尔可夫可行性和标准动态规划正则性下,PA贝尔曼选择器诱导出最优全历史策略。在附加矩形性和穷举性条件下,将偏好记忆重新参数化为不同的信念和品味坐标,得到分离表示和贝尔曼递归。还给出示例分类以说明框架范围。
英文摘要
In finite horizon dynamic programming with history-dependent preferences, the relevant state may be the entire realized history, even when the physical state is Markov. This paper develops a behavioral state-reduction theory for such Markov decision processes. Under behavioral axioms and a certainty-equivalent richness condition, the full-history problem admits a recursive representation composed of time and risk aggregators. We then derive a canonical preference-augmented (PA) state by quotienting histories that have the same current physical Markov state, are indifferent under every common continuation plan, and remain equivalent after every common one-step extension. This canonical PA state is minimal among reachable recursive factorizations of the underlying preferences. Under Markov feasibility and standard dynamic-programming regularity, a PA Bellman selector induces an optimal full-history policy. With additional rectangularity and exhaustiveness conditions, we reparameterize the preference memory into distinct belief and taste coordinates, and obtain a separated representation and Bellman recursion. We give a taxonomy of examples to illustrate the scope of our framework.