arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11433cs.RO

动态环境中强化学习的安全感知技能适应

Safety-aware Skill Adaptation for Reinforcement Learning in Dynamic Environments

  • University of Technology Sydney(悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

A K M Nadimul Haque, Sheila Sutjipto, Marc G. Carmichael, Teresa Vidal-Calleja

AI总结:

针对动态环境中的机器人技能适应,提出Dist-GPRL框架,利用高斯过程参数化局部窗口更新轨迹,结合安全子空间先验和距离场奖励引导探索,在模拟和真实任务中实现更高成功率、更低碰撞频率和更稳定学习。

AI中文摘要:

基于强化学习的技能适应框架通常需要限制性假设以维持稳定性,例如固定观测或严格控制探索计划。然而,在杂乱且动态的环境中,不受限制的探索可能导致不安全行为和训练不稳定,特别是当任务相关观测位于障碍物附近或涉及移动物体时。在本工作中,我们提出了Dist-GPRL,一种用于结构化机器人技能适应的距离感知和安全引导的强化学习框架。基于高斯过程(GP)的技能参数化,我们的框架顺序地适应稀疏轨迹路点的重叠局部窗口,而不是在每个策略步骤修改完整技能。原始策略输出通过GP协方差结构相关联,产生时间上连贯的轨迹更新,同时减少了与全局轨迹适应相关的动作空间和信用分配困难。安全性通过两种互补的引导形式纳入。源自Hausdorff近似规划器(HAP)的安全子空间先验将策略探索偏向可行区域,而动态更新的距离场间隙和梯度奖励提供局部障碍感知。轨迹-运动学相似性正则化器进一步在适应过程中保持演示的速度和加速度特征。我们在模拟中的两个动态物体操作任务上评估了该框架,并将学习到的策略迁移到真实机器人执行。实验结果表明,与基线相比,任务成功率更高,碰撞频率更低,学习更稳定,同时保持了演示技能的动力学特征。

英文摘要:

Skill adaptation frameworks based on reinforcement learning often require restrictive assumptions to maintain stability, such as fixed observations or tightly controlled exploration schedules. In cluttered and dynamic environments, however, unrestricted exploration can lead to unsafe behaviour and unstable learning, particularly when task-relevant observations lie near obstacles or involve moving objects. In this work, we present Dist-GPRL, a distance-aware and safety-guided reinforcement learning framework for structured robot skill adaptation. Building upon Gaussian Process (GP)-based skill parameterisation, our framework sequentially adapts overlapping local windows of sparse trajectory via-points rather than modifying the complete skill at every policy step. Raw policy outputs are correlated through the GP covariance structure, producing temporally coherent trajectory updates while reducing the action-space and credit-assignment difficulties associated with global trajectory adaptation. Safety is incorporated through two complementary forms of guidance. A safe-subspace prior derived from the Hausdorff Approximation Planner (HAP) biases policy exploration toward feasible regions, while dynamically updated distance field clearance and gradient rewards provide local obstacle awareness. A trajectory-kinematics similarity regulariser further preserves the demonstrated velocity and acceleration characteristics during adaptation. We evaluate the framework on two dynamic object-manipulation tasks in simulation and transfer the learned policy to real-world robot execution. Experimental results demonstrate higher task success, lower collision frequency, and more stable learning than the baselines, while preserving the kinematic characteristics of the demonstrated skill.

补充信息

↑