SegBench-GC:多步离线目标条件强化学习中的分割不变性测试
SegBench-GC: Testing Segmentation Invariance in Multi-Step Offline Goal-Conditioned Reinforcement Learning
中文总结 AI 辅助
本文提出SegBench-GC基准,通过受控测试发现离线目标条件强化学习中,管理性分割会显著改变多步任务性能,且CVT处理方式可缓解该问题。
中文摘要 AI 辅助
离线目标条件强化学习(GCRL)常利用轨迹结构进行未来目标采样和多步目标设定,但记录的轨迹可能因管理原因被分割,这种分割与终止状态无关。本文提出SegBench-GC,这是一种针对分割不变性的受控压力测试,在保持转换、源轨迹、目标采样、优化设置和评估固定的前提下,仅改变人工备份边界以及这些边界是否保留延续价值。延续有效目标(CVT)提供与分割一致的控制:奖励积累在人工分割处停止,但目标从其存储的后续状态引导。在匹配数量的PointMaze研究中,设置35000个人工分割、三种分割实现方式和三个优化种子,最终每任务50回合的成功率为:未分割时50.5%,采用CVT时39.1%,将相同分割视为吸收状态时19.1%;跨分割实现方式,朴素平均成功率范围为4.8%至31.9%。来自Decoupled Q-Chunking代码库的独立已发表n步基线(n=25)在Puzzle-4x5上显示相同失败情况:三个优化种子下未分割时47.2%,采用CVT时58.5%,朴素处理时0.27%。目标级诊断验证了分析目标差异的数值精度,学习到的评论者诊断显示,朴素处理下存在大幅乐观偏移,而CVT仍与未分割评论者大致对齐。CVT应用标准延续引导而非新的Bellman规则,本文的贡献在于提供了受控基准、失败隔离方法以及跨学习器的证据,证明管理性分割可显著改变多步离线GCRL的性能。
英文摘要
Offline goal-conditioned reinforcement learning (GCRL) often uses trajectory structure for future-goal sampling and multi-step targets, yet logged trajectories may be partitioned for administrative reasons that do not correspond to termination. We introduce SegBench-GC, a controlled stress test of segmentation invariance that holds transitions, source trajectories, goal sampling, optimization settings, and evaluation fixed while varying only artificial backup boundaries and whether those boundaries retain continuation value. Continuation-valid targets (CVT) provide the segmentation-consistent control: reward accumulation stops at an artificial cut, but the target bootstraps from its stored successor. In a matched-count PointMaze study with 35,000 artificial cuts, three segmentation realizations, and three optimization seeds, final 50-episode-per-task success is 50.5% uncut, 39.1% with CVT, and 19.1% when the same cuts are treated as absorbing; across segmentation realizations, naive mean success ranges from 4.8% to 31.9%. An independent published n-step baseline (n=25) from the Decoupled Q-Chunking codebase shows the same failure on Puzzle-4x5: 47.2% uncut, 58.5% CVT, and 0.27% naive across three optimization seeds. A target-level diagnostic verifies the analytic target difference to numerical precision, and learned-critic diagnostics show a large optimistic shift under naive handling while CVT remains approximately aligned with the uncut critic. CVT applies standard continuation bootstrapping rather than a new Bellman rule; the contribution is the controlled benchmark, failure isolation, and cross-learner evidence that administrative segmentation can materially change multi-step offline GCRL.