SCoCaT: 用于航天器对接的成功条件约束强化学习
SCoCaT: Success Conditioned Constrained Reinforcement Learning for Spacecraft Docking
浏览论文内容
中文总结 AI 辅助
针对航天器对接中基于终止的约束强化学习在终端导航任务中因目标靠近约束激活区域导致任务完成率低的问题,提出通过辅助价值评论家添加密集成功信号的方法,提高任务完成率并保持安全合规性。
中文摘要 AI 辅助
基于终止的约束强化学习对于安全关键的机器人部署具有吸引力:它避免了推理时的在线优化,通过每个约束的单一标量轻松扩展到多个约束,并且比常用的拉格朗日方法更易于实现。该方法不是通过累加成本惩罚来对违规行为定价,而是通过缩短每次违规的有效时间范围,使违规行为在结构上无利可图。我们识别出该方法类在终端导航任务中的一个结构性失效模式:在满足沿最终接近路径收紧的安全约束的同时,达到精确的目标配置。当目标位于接近约束激活区域的内部时,生存加权目标使得停留在目标区域之外严格优于进入目标区域,从而产生高约束合规性但低任务完成率。我们形式化了这一病理,并表明对现成的强化学习算法(如PPO)进行最小增强即可解决这种“可行性崩溃”。我们通过实验证明,通过辅助价值评论家添加密集的每步成功信号,可以提高任务完成率,同时保持安全关键的约束合规性。在两个代表性航天器平台上的验证支持了这些发现的普遍性:一个覆盖运行邻近操作质量和自由度范围的6U立方星,以及我们实验室中用于零样本模拟到现实迁移的浮动平台测试台。
英文摘要
Termination-based constrained reinforcement learning is attractive for safety-critical robotic deployments: it avoids online optimization at inference, scales easily to many constraints via a single scalar per constraint, and is simpler to implement than commonly used Lagrangian methods. Instead of pricing violations through summed cost penalties, this approach makes violations structurally unprofitable by shortening the effective horizon for each violation. We identify a structural failure mode of this method class on terminal-navigation tasks: reaching a precise goal configuration while satisfying safety constraints that tighten along the final approach. When the goal sits inside the region close to where the constraints become active, the survival-weighted objective makes dwelling outside the goal region strictly preferable to entering, producing high constraint compliance with low task completion. We formalize this pathology and show that a minimal augmentation to off-the-shelf RL algorithms like PPO resolves this ``feasibility collapse''. We empirically demonstrate that adding a dense per-step success signal via an auxiliary value critic improves the task completion rate while maintaining safety-critical constraint compliance. Validation across two representative spacecraft platforms: a 6U-CubeSat spanning the mass and degree-of-freedom envelope of operational proximity operations, and a floating platform testbed for zero-shot sim-to-real transfer in our laboratory, supports the generality of these findings.
发表机构
- SnT - Interdisciplinary Centre for Security, Reliability and Trust(安全、可靠与信任跨学科中心)
- University of Luxembourg(卢森堡大学)
机构由 AI 辅助整理,请以论文原文为准。