混合刚体-气动操纵器的耦合状态空间建模、控制与策略蒸馏
Coupled State-Space Modelling, Control, and Policy Distillation for Hybrid Rigid-Pneumatic Manipulators
查看机构详情
- Indian Institute of Technology Madras(印度理工学院马德拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文为混合刚体-气动操纵器建立耦合模型,证明解耦控制的代价,并将MPC蒸馏为小型神经策略,实现零碰撞下93-94%的稳定率且满足实时性要求。
中文摘要 AI 辅助
混合操纵器将电机驱动的刚性关节与压力驱动的折纸段相结合。已发表的此类机械臂采用解耦的每自由度控制回路进行控制,而这种近似所带来的代价尚未被量化,因为用于测量该代价的耦合模型尚未建立。本文针对由N个交替排列的旋转关节和Kresling折纸段组成的链推导了这样一个模型,其中包括气动腔室动力学和折痕迟滞。利用该模型,我们直接测量了耦合效应,并表明其强度随关节不同而变化,且解耦控制在强耦合关节上恰好失效,而在几乎解耦的关节上仍具有竞争力。基于耦合模型的控制器在较低扭矩下比解耦PID基线跟踪精度高2.5倍。然而,模型预测控制器(MPC)对于实时应用而言速度过慢,而无模型强化学习在严格的稳定指标下停滞在远低于可接受的成功率水平。因此,我们通过行为克隆和DAgger将MPC蒸馏为一个小型神经策略。蒸馏后的策略在零碰撞的情况下实现了93-94%的目标稳定率,与教师策略相差几个百分点,并在5毫秒的控制步长内运行,而MPC则无法做到。在教师策略本身失败的情况下,我们将失败追溯到与波纹管轻阻尼模式相关的极限环,并通过选择可在低压下保持的目标姿态来消除该问题。
英文摘要
Hybrid manipulators combine motorized rigid joints with pressure-actuated origami segments. Published arms of this kind are controlled with decoupled per-DOF loops, and the cost of this approximation has not been quantified, because the coupled model needed to measure it has not been built. This paper derives such a model for a chain of $N$ alternating revolute joints and Kresling origami segments, including pneumatic chamber dynamics and crease hysteresis. Using the model, we measure the coupling directly and show that its strength varies joint by joint, and that decoupled control loses precisely on the strongly coupled joints while remaining competitive on the one nearly decoupled joint. Coupled model-based controllers track $2.5\times$ tighter than a decoupled PID baseline at lower torque. However, the model predictive controller (MPC) is too slow for real time, and model-free reinforcement learning stalls far below acceptable success rates on a strict settling metric. We therefore distill the MPC into a small neural policy with behavior cloning and DAgger. The distilled policy settles 93-94$\%$ of goals with zero collisions, within a few points of its teacher, and runs inside the 5 ms control step where the MPC does not. Where the teacher itself fails, we trace the failure to a limit cycle with the bellows' lightly damped mode, and we remove it by selecting goal postures holdable at low pressure.