完整覆盖决策过程中通过稳定商实现的最小马尔可夫化
Minimal Markovization via Stable Quotients in Holonomy-Cover Decision Processes
浏览论文内容
中文总结 AI 辅助
本文针对结构化POMDP类刻画最小马尔可夫充分统计量,构造稳定商实现精确有限马尔可夫状态,提出完整记忆强化学习方法,实验验证其压缩与决策性能优于基线。
中文摘要 AI 辅助
在部分可观测环境中行动的智能体必须保留可递归更新的历史统计量以恢复马尔可夫性质,但这类最小统计量通常未知。本文针对完整覆盖决策过程(Holonomy-Cover Decision Processes)——一类可见动力学为马尔可夫且每个已实现的可见转换对隐藏模式应用固定置换的结构化部分可观测马尔可夫决策过程(POMDP)——刻画了其最小马尔可夫充分统计量。具体而言,本文构造了稳定商(stable quotient),这是保留一步奖励和商后继的最粗粒度观测级抽象,并证明当前观测与稳定类的对构成精确有限马尔可夫状态。当当前类被正确初始化时,精确类跟踪恰好需要最小内存符号:在可达性和最大化观测下的成对决策分离条件下,不存在任意有限内存控制器能使用更少符号。在可重置诊断下,近原型类推理具有指数衰减误差,且校准后重启约简将有限马尔可夫决策过程(MDP)保证传递到恢复状态。这些结果支持完整记忆强化学习(Holonomy Memory Reinforcement Learning):它通过当前稳定类表示内存,通过有序边传输更新,在诊断可用时识别局部类坐标,并在同步后应用标准有限MDP强化学习主干。实验从原始状态恢复到商状态的精确压缩,用三个决策时间内存状态实现完美配对顺序准确率,匹配商神谕且优于非神谕基线。
英文摘要
An agent acting under partial observability must retain a recursively updateable statistic of history that restores the Markov property, but the smallest such statistic is generally unknown. We characterize this minimal Markov sufficient statistic for holonomy-cover decision processes, a structured POMDP class in which the visible dynamics are Markov and every realized visible transition applies a fixed permutation to a hidden mode. In particular, we construct the stable quotient, the coarsest observation-wise abstraction preserving one-step rewards and quotient successors, and prove that the pair of the current observation and stable class forms an exact finite Markov state. When the current class is correctly initialized, exact class tracking requires exactly the minimal memory symbols, in the sense that under reachability and pairwise decision separation at a maximizing observation, no arbitrary finite-memory controller can use fewer. Under resettable diagnostics, nearest-prototype class inference has exponentially decaying error, and a calibrate-then-restart reduction transfers finite-MDP guarantees to the recovered state. The results enable \emph{Holonomy Memory Reinforcement Learning}. It represents memory by the current stable class, updates it through ordered edge transports, identifies local class coordinates when diagnostics are available, and applies a standard finite-MDP RL backbone after synchronization. Experiments recover an exact compression from raw states to quotient states and achieve perfect paired-order accuracy with three decision-time memory states, matching the quotient oracle and outperforming the non-oracle baselines.