学习速度自适应的髋部外骨骼控制策略:基于仿真到现实强化学习
Learning a Speed-adaptive Hip Exoskeleton Control Policy Via Sim-to-real Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
提出结合仿真到现实强化学习与在线偏好学习的框架,解耦辅助时机与幅度优化,实现速度自适应的个性化髋部外骨骼辅助,减少真实世界评估次数。
中文摘要 AI 辅助
在不同行走速度下提供个性化外骨骼辅助仍然具有挑战性。现有的在线优化方法样本效率低下,需要大量的人机交互(HIL)评估来优化整个辅助力矩曲线。仿真到现实强化学习(RL)提供了一种有前景的替代方案,但无法直接考虑个体用户偏好。我们提出一个将仿真到现实强化学习与在线偏好学习相结合的框架,用于个性化外骨骼辅助。具体而言,辅助时机通过在仿真中训练强化学习策略与人体肌肉骨骼模型在不同行走速度下学习。学习到的策略随后被蒸馏并部署在物理髋部外骨骼上,利用机载感官观测。基于高斯过程的偏好学习通过成对用户比较进一步个性化辅助幅度。通过将仿真中的辅助时机学习与真实世界实验中的幅度优化解耦,我们的框架大幅减少了在线优化空间。人体受试者实验表明,在更少的真实世界评估下,能够高效识别不同行走速度下的个性化辅助力矩曲线。
英文摘要
Providing personalized exoskeleton assistance across varying walking speeds remains challenging. Existing online optimization methods are sample-inefficient, requiring extensive human-in-the-loop (HIL) evaluations to optimize the entire assistive torque profile. Sim-to-real reinforcement learning (RL) offers a promising alternative but cannot directly account for individual user preferences. We propose a framework integrating sim-to-real RL with online preference learning for personalized exoskeleton assistance. Specifically, assistance timing is learned in simulation by training RL policies with human musculoskeletal models across varying walking speeds. The learned policies are then distilled and deployed on a physical hip exoskeleton using onboard sensory observations. Gaussian-process-based preference learning further personalizes the assistance magnitude through pairwise user comparisons. By decoupling assistance timing learning in simulation from magnitude optimization in real-world experiments, our framework substantially reduces the online optimization space. Human-subject experiments demonstrate efficient identification of personalized assistive torque profiles across varying walking speeds with fewer real-world evaluations.
发表机构
- Lingnan University(岭南大学)
- Southern University of Science and Technology(南方科技大学)
机构由 AI 辅助整理,请以论文原文为准。