arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

揭示覆盖之后:历史相关日志下指数级POMDP离策略评估下界

Revealing After Overwriting: An Exponential POMDP OPE Lower Bound under History-Dependent Logging

Youyu Luo, Pengzhan Zhou, Zhida Qin, Jia Wang, Zuotao Fu, Yu Liu, Chao Chen

arXiv 2610.05063首次发表:更新:

发表机构

Chongqing University; Beijing Institute of Technology; The Hong Kong Polytechnic University(重庆大学; 北京理工大学; 香港理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对历史相关日志下的POMDP离策略评估,构造指数级下界并给出有限类保证,结合读取预算与校准实现与水平无关的样本复杂度。

AI 中文摘要

多步揭示可以在无记忆日志下使离策略评估变得可处理。在历史相关日志下,状态可解码性与目标相关证据可能分离。对于每个水平$H\ge3$,我们构造两个完全可实现的部分可观测马尔可夫决策过程(POMDP),具有四个动作、每层至多四个状态、一个已知记录器和一个无记忆目标。动作重叠、历史覆盖和仅观测揭示保持有界,与$H$无关,但目标值相差$1/2$,且记录定律之间的KL散度为$\Theta(4^{-(H-1)})$,迫使样本复杂度呈指数级。记录器记忆使状态可区分,而重置则抹除了目标所保留的模型区分证据。一个单独的构造在常见、已知的仅观测揭示算子下保留此障碍。在动作和历史覆盖下,我们给出一个有限类别的离策略评估保证,使用在每段历史中有效的常见可观测值表示。样本界多项式依赖于它们的二阶矩成本。在常见算子构造中,相同的值方向具有恒定的边际解码成本,但指数级的历史条件成本。最后,在一个固定的四动作连续体上,我们推导出匹配的被动和预算读取速率。在已知信道和单位读取成本下,早期读取是最优的。在未知传感器偏差下,仅早期读取仍保持指数级成本。当两种读取类型在每条轨迹上获得固定的正期望预算时,将它们与重置后校准相结合,样本复杂度与$H$无关。

英文摘要

Multi-step revealing can make off-policy evaluation tractable under memoryless logging. With history-dependent logging, state decodability and target-relevant evidence can separate. For every horizon $H\ge3$, we construct two exactly realizable POMDPs with four actions, at most four states per layer, a known logger, and a memoryless target. Action overlap, history coverage, and observation-only revealing remain bounded independently of $H$, yet the target values differ by $1/2$ and the KL divergence between the logged laws is $Θ(4^{-(H-1)})$, forcing exponential sample complexity. Logger memory makes states distinguishable, while reset erases the model-distinguishing evidence preserved by the target. A separate construction retains this barrier with common, known observation-only revealing operators. Under action and history coverage, we give a finite-class OPE guarantee using common observable value representations that remain valid at every history. The sample bound depends polynomially on their second-moment cost. In the common-operator construction, the same value direction has constant marginal decoding cost but exponential history-conditioned cost. Finally, on a fixed four-action continuum, we derive matching passive and budgeted readout rates. With one known channel and unit read cost, early reads are optimal. With unknown sensor bias, early reads alone remain exponentially costly. Combining them with post-reset calibration gives sample complexity independent of $H$ when both read types receive fixed positive expected budgets per trajectory.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑