面向机器人机械臂扰动补偿的稳定性感知残差强化学习框架
Stability-aware Residual Reinforcement Learning Framework for Robotic Manipulator Disturbance Compensation
- Hanyang University(汉阳大学)
- Kyonggi University(京畿大学)
- Kookmin University(国民大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对机械臂受扰动影响的问题,提出将解析观测器与强化学习策略结合的残差框架,并基于ISS分析保证稳定性,在真实硬件上实现最高38%的跟踪误差降低。
中文摘要 AI 辅助
尽管传统控制器和扰动观测器(DOBs)是机械臂精确跟踪的标准方法,但它们仍面临参数不确定性、非线性摩擦和复合扰动等问题。本研究提出了一种残差强化学习DOB框架,将解析观测器与强化学习(RL)策略相结合。确定性基线在可靠区域内运行,而RL策略则明确针对模型无法捕捉的残差。为使这种补偿具有扰动感知能力,估计器网络将观测历史与特权扰动上下文对齐,按扰动类型组织潜在空间,从而在扰动转换期间实现快速适应。为保证稳定性,我们基于输入到状态稳定性(ISS)分析推导并强制执行了RL策略的状态相关动作界限,使得闭环能够证明地将跟踪误差限制在认证包络内,无论策略输出如何。在六自由度机械臂上的实验表明,扰动估计和跟踪性能均得到持续改善,包括在零样本仿真到现实迁移下真实硬件上跟踪误差降低27.8%,以及在训练期间未观测到的基座振动扰动下误差降低38.0%。
英文摘要
Although conventional controllers and disturbance observers (DOBs) are the standard for precision tracking in manipulators, they suffer from parameter uncertainty, nonlinear friction, and compound disturbances. This study proposes a residual reinforcement learning DOB framework that pairs an analytical observer with an RL policy. The deterministic baseline operates within a reliable region, whereas the RL policy explicitly targets the residuals that the model cannot capture. To make this compensation disturbance-aware, an estimator network aligns the observation history with a privileged disturbance context, organizing the latent space by disturbance regime and enabling rapid adaptation across disturbance transitions. To guarantee stability, we derived and enforced a state-dependent action bound on the RL policy from an input-to-state stability (ISS) analysis such that the closed loop provably confines the tracking error to a certified envelope for arbitrary policy outputs. Experiments on a 6-DOF manipulator demonstrated consistent improvements in disturbance estimation and tracking, including a 27.8% tracking-error reduction on real hardware under zero-shot sim-to-real transfer and a 38.0% reduction under a base-vibration disturbance that was not observed during training.