发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无重置强化学习在不可逆环境中易陷入不可恢复状态的问题,提出REVERSAL-BENCH基准,通过可逆性参数和重置预言机揭示可逆性悬崖,并验证安全盾牌干预的有效性。
AI 中文摘要
自主强化学习的核心目标是在没有外部重置的情况下持续进行策略训练。然而,现有范式在很大程度上依赖于环境的内在可逆性,而这一属性在现实世界的操作中并不存在,例如将物体推下桌子或洒出颗粒状物质等事件无法被撤销。我们引入了REVERSAL-BENCH,这是一个通过连续参数$\rho \in [0, 1]$控制可逆性并提供重置预言机(reset oracle)的基准测试平台,该预言机是一种用于在五个物理引擎中的八种操作场景下测试状态可恢复性的真值验证机制。评估涵盖广泛的策略架构,包括标准演员-评论家算法、安全强化学习以及专门的无重置框架,结果揭示了一个尖锐的可逆性悬崖:随着$\rho$的增加,无重置智能体持续被吸收到不可恢复状态中,而回合制智能体则保持稳定的学习。我们在自主无重置基线和约束强化学习中均观察到这种失败模式。由于无重置智能体缺乏外部重置,任何进入不可恢复状态的转移都会导致永久吸收,使智能体被困住,进一步学习随之停止。我们证明,这种吸收现象在学习到的操作策略下的完整物理模拟中持续存在。通过与几何上相同的可逆对应物进行对比评估,我们确认这种崩溃是由不可逆性而非障碍物复杂性因果驱动的。我们发布了基准测试套件、一个带有可恢复性标注的大规模多模拟器数据集以及重置预言机。我们还评估了一种在不可逆失败发生前进行干预的安全盾牌,结果表明,虽然可恢复性可以被准确预测,但主动恢复主要仅在智能体能够物理上避开陷阱时才能成功。
英文摘要
A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying environmental reversibility, a property absent in real world manipulation, where events such as pushing objects off tables or spilling granular substances cannot be undone. We introduce REVERSAL-BENCH, a benchmark that controls reversibility via a continuous parameter $ρ\in [0, 1]$ and provides a reset oracle, a ground-truth verification mechanism to test state recoverability across eight manipulation settings in five physics engines. Evaluating a broad spectrum of policy architectures, including standard actor-critic algorithms, safe RL, and specialized reset-free frameworks, reveals a sharp reversibility cliff: reset-free agents are consistently absorbed into irrecoverable states as $ρ$ increases, whereas episodic agents maintain steady learning. We see this failure mode across autonomous reset-free baselines and constrained RL. Because reset-free agents lack external resets, any transition into an irrecoverable state results in permanent absorption, leaving the agent trapped where further learning halts. We show that this absorption phenomenon persists in full physics simulations under learned manipulation policies. By evaluating against geometrically identical reversible counterparts, we confirm that this breakdown is causally driven by irreversibility rather than obstacle complexity. We release the benchmark suite, a large multi-simulator dataset labeled with recoverability and a reset oracle. We also evaluate a safety shield that intervenes before irreversible failures occur, showing that while recoverability can be predicted accurately, active recovery primarily succeeds only when the agent can physically steer clear of the trap
Comments8 pages, 5 figures, 2 tables