SwingRL:基于世界模型预测的自适应观测强化学习用于缆索悬挂吊装控制
SwingRL: Adaptive Observation Reinforcement Learning with World-Model Prediction for Cable-Suspended Hoisting Control
AI总结:
SwingRL提出结合年龄感知世界模型、经典防摆先验和循环残差策略的强化学习框架,解决缆索悬挂吊装中的过时反馈和不确定动力学问题,在困难设置下显著优于基线,并无需微调即在物理台实现90%成功率。
AI中文摘要:
缆索悬挂吊装被广泛用于移动无法由刚性拾放系统方便处理的沉重或庞大载荷,例如在起重机辅助施工中。使用柔性缆索的机器人吊装具有挑战性,因为载荷运动是欠驱动的,外部扰动变化多端,且延迟或丢失的视觉观测可能使感知到的载荷状态在控制执行时变得过时。这些影响在将悬挂载荷的插座精确插入钢筋销钉时尤为关键,这是施工环境中非常常见的任务。我们提出了SwingRL,一个残差强化学习(RL)框架,结合了年龄感知世界模型、经典防摆先验和循环残差策略,以解决两个耦合问题:过时反馈和不确定动力学。世界模型利用已执行的命令将最新接收到的载荷观测传播到当前控制步骤,在延迟和有损感知下提供时间对齐的状态估计。先验提供标称跟踪和防摆阻尼。残差策略学习对先验的有界修正,而非完整控制律,补偿系统参数变化、外部扰动和剩余的状态估计误差。我们将SwingRL与经典和基于学习的基线在涵盖系统参数变化、风扰动、退化感知和强阵风的累积难度阶梯上进行了评估。在最困难的设置下,SwingRL实现了69.5%的严格成功率和77.3%的广泛成功率,分别超过所有基线至少60个百分点。世界模型消融研究支持了时间对齐状态估计在观测损失增加时维持插入成功的作用。最后,在没有真实机器人微调的情况下,SwingRL在物理实验台上实现了90%的成功率。
英文摘要:
Cable-suspended hoisting is widely used to move heavy or bulky payloads that cannot be handled conveniently by rigid pick-and-place systems, for example in crane-assisted construction. Robotic hoisting using flexible cables is challenging because payload motion is underactuated, external disturbances vary, and delayed or lost visual observations can make the perceived payload state stale at control execution. These effects are particularly critical during precise insertion of a suspended payload's sockets onto rebar pins, which is a very common task in construction environments. We present SwingRL, a residual reinforcement-learning (RL) framework that combines an age-aware world model, a classical anti-swing prior, and a recurrent residual policy to address two coupled problems: stale feedback and uncertain dynamics. The world model propagates the newest received payload observation to the current control step using the executed commands, providing a time-aligned state estimate under delayed and lossy sensing. The prior supplies nominal tracking and swing damping. The residual policy learns bounded corrections to the prior rather than the complete control law, compensating for system-parameter variation, external disturbances, and remaining state-estimation errors. We evaluate SwingRL against classical and learning-based baselines across a cumulative difficulty ladder covering system-parameter variation, wind disturbance, degraded sensing, and strong gusts. Under the most difficult setting, SwingRL achieves 69.5% strict and 77.3% broad success, exceeding all baselines by at least 60 percentage points, respectively. World-model ablations support the role of time-aligned state estimation in maintaining insertion success as observation loss increases. Finally, without real-robot fine-tuning, SwingRL achieves 90% success on the physical rig.