发表机构
The Pennsylvania State University; University of Notre Dame(宾夕法尼亚州立大学; 圣母大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种针对表格强化学习的精确快速批量模拟框架,通过聚合马尔可夫流和前向采样将计算复杂度从O(m)降至O(1)或O(log m),适用于多种RL设置。
AI 中文摘要
模拟是强化学习(RL)中的一个基本计算原语,然而传统的模拟即使在下游过程仅使用聚合统计量时,也会显式地生成单个轨迹。为了解决这一问题,我们为有限时域表格马尔可夫决策过程开发了一个精确的快速模拟框架。该框架具有两种互补模式。在直接批量模拟中,一个批次由其聚合马尔可夫流表示。在具有足够的并行模拟资源时,该流可通过轨迹聚合获得;当此类模拟不可用或代价高昂,但初始状态和转移分布可直接访问时,我们则通过前向马尔可夫流采样生成同分布的流,而无需具体化单个轨迹。后者将模拟器侧对批次大小$m$的计算依赖性从$O(m)$降低到$O(1)$。在自适应批量模拟中,当批次长度由数据依赖条件决定时,精确的多元超几何分裂递归地细化候选马尔可夫流,同时保持条件分布,将成本对$m$的依赖性从$O(m)$降低到$O(\log m)$。这些模式共同通过尽可能保持轨迹聚合、仅在需要定位数据依赖边界时细化流来加速模拟。该框架广泛适用于基于模拟器、离线和在线的批量或阶段式RL,并通过每种设置中的代表性算法加以说明。
英文摘要
Simulation is a fundamental computational primitive in reinforcement learning (RL), yet conventional simulation explicitly generates individual trajectories even when downstream procedures use only aggregate statistics. To address this, we develop an exact fast-simulation framework for finite-horizon tabular Markov decision processes. Our framework has two complementary modes. In direct batch simulation, a batch is represented by its aggregate Markov flow. With sufficient parallel simulation resources, this flow can be obtained by trajectory aggregation; when such simulation is unavailable or costly but the initial state and transition distributions are directly accessible, we instead generate an identically distributed flow through forward Markov-flow sampling without materializing individual trajectories. The latter reduces the simulator-side computational dependence on batch size $m$ from $O(m)$ to $O(1)$. In adaptive batch simulation, when batch length is determined by a data-dependent condition, exact multivariate-hypergeometric splitting recursively refines a candidate Markov flow while preserving the conditional law, reducing the cost dependence on $m$ from $O(m)$ to $O(\log m)$. Together, these modes accelerate simulation by keeping trajectories aggregated whenever possible and refining flows only when required to locate data-dependent boundaries. The framework applies broadly across simulator-based, offline, and online batch or stage-based RL, as illustrated with representative algorithms from each setting.