arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17946cs.RO

反馈调制的谐波策略用于四足 locomotion

Feedback-Modulated Harmonic Policies for Quadruped Locomotion

  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Yixuan Jia, Steven Roche, Jonathan P. How

AI总结:

本文提出反馈调制谐波策略,将四足运动关节轨迹表示为命令条件傅里叶级数,通过上下文网络生成系数并在线调整,在仿真和 Unitree Go2 上验证了周期性结构与高性能。

AI中文摘要:

学习到的四足 locomotion 策略通常将观测直接映射到关节级动作,使得 locomotion 的周期性结构在策略中隐式存在。我们研究了一种替代表示,其中每个关节轨迹被表示为命令条件傅里叶级数,并根据机器人状态的反馈在线修改。一个上下文网络生成傅里叶系数和每步反馈网络的权重,该反馈网络的输出在执行过程中调整关节偏移、谐波增益、频率和相位。在仿真中,我们检查了这种显式频率结构以及直接输出关节目标的 MLP 策略的隐藏激活。谐波波形随命令速度改变频率和形状。对选定 MLP 轨迹的动态模式分解揭示了在脚高度振荡频率及其二次谐波附近的优势激活模式,表明在没有显式傅里叶生成器的情况下存在周期性结构。在 Unitree Go2 上,仿真训练的谐波控制器在单独试验中记录了临时机载估计峰值速度 3.67 米每秒,并承载额外负载高达 5.883 千克。

英文摘要:

Learned quadruped locomotion policies commonly map observations directly to joint-level actions, leaving the periodic structure of locomotion implicit in the policy. We investigate an alternative representation in which each joint trajectory is expressed as a command-conditioned Fourier series and modified online using feedback from the robot state. A context network generates the Fourier coefficients and the weights of a per-step feedback network, whose outputs adjust joint offsets, harmonic gains, frequency, and phase during execution. In simulation, we examine this explicit frequency structure alongside the hidden activations of an MLP policy that directly outputs joint targets. The harmonic waveforms change frequency and shape with commanded speed. Dynamic mode decomposition of selected MLP rollouts reveals dominant activation modes near the foot-height oscillation frequency and its second harmonic, showing periodic structure without an explicit Fourier generator. On a Unitree Go2, the simulation-trained harmonic controller records a provisional onboard-estimated peak speed of 3.67 meter per second and carries added loads up to 5.883 kilogram in separate trials.

补充信息

↑