非马尔可夫决策过程中的精确可区分性
Exact Distinguishability in Non-Markovian Decision Processes
浏览论文内容
中文总结 AI 辅助
本研究针对非马尔可夫决策过程,提出PEC算法在线性时间内判定模型可区分性,并证明等价模型后验几率不变,实验表明该算法能恢复失效的可区分性假设。
中文摘要 AI 辅助
非马尔可夫环境通常被建模为正则决策过程(RDPs),其中动力学通过有限自动机依赖于交互历史。现有的离线保证依赖于对行为策略的可区分性假设,但未提供验证该假设的方法。当该假设被违反时,不同的模型可能同样好地解释数据。我们研究了在固定行为策略下收集的数据何时能区分两个候选RDP。我们证明了在观测上等价的候选者之间的后验几率在每一个样本量下都保持等于先验几率,即使策略访问了每个自动机状态,并在Lean 4中正式验证了这两个结果。然后,我们精确刻画了这一等价性,并推导出PEC算法,该算法在乘积自动机大小的线性时间内判定等价性。先前工作的可区分性假设在我们的四个测试环境中的三个上失败,而PEC识别的实验在每种情况下都恢复了该假设。
英文摘要
Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton. Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it. When the assumption is violated, distinct models may explain the data equally well. We study when data collected under a fixed behaviour policy can distinguish two candidate RDPs. We prove that the posterior odds between observationally equivalent candidates remain equal to the prior odds at every sample size, even when the policy visits every automaton state, and verify both results formally in Lean 4. We then characterize this equivalence exactly and derive PEC, an algorithm that decides it in time linear in the size of the product automaton. The distinguishability assumption of prior work fails on three of our four test environments, and the experiment identified by PEC restores it in each case.
发表机构
- Nirma University(尼尔玛大学)
- Google(谷歌)
机构由 AI 辅助整理,请以论文原文为准。