arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StableMimic:人形机器人运动跟踪的类人平滑恢复——超越跟踪分布学习结构化跌倒后行为

StableMimic: Smooth Human-Like Recovery for Humanoid Motion Tracking - Learning Beyond the Tracking Distribution for Structured Post-Fall Behavior

Weihao Wu, Ming Huang, Ruofei Liu, Jinglei Nie, Shuxiang Guo, Chunying Li

arXiv 2608.02385首次发表:更新:

发表机构

Southern University of Science and Technology (SUSTech)(南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出StableMimic,一种在标称跟踪分布外训练的人形机器人统一跟踪器,通过专用专家与本体感知门控融合动作,在LAFAN1数据集及真实Unitree G1机器人上实现跌倒后100%恢复,跟踪与恢复性能均优于对比方法,提升了交互安全性。

AI 中文摘要

人形机器人运动跟踪器在学习到的跟踪分布内表现可靠,但跌倒会使机器人进入低高度、接触密集的状态,此时前进指令暂时无法触及。仅跟踪的策略可能会追逐不可行的参考,产生快速、大幅的肢体修正,增加机器人及其周围环境的风险。我们提出StableMimic,这是一个在标称跟踪分布之外训练的统一跟踪器。围绕多个人类起身参考的扰动重置,暴露了俯卧、仰卧、失衡和中间地面接触状态,塑造了将机器人返回可跟踪区域的结构化恢复。由于跟踪和恢复占据明显不同的状态-动作分布,StableMimic对每种模式使用专用专家,以及一个本体感知门控,该门控持续融合它们的动作。隐藏的后继状态目标教授类人参考形状的恢复,而不向部署的Actor暴露参考身份或阶段;部署不需要起身参考、恢复指令、轨迹检索或外部策略切换。在完整的重定向LAFAN1舞蹈子集上,StableMimic在五种方法中的所有四个跟踪指标上实现了最低误差。在每种方法的100次匹配推至跌倒试验中,它100/100恢复,并在7项跌倒后运动和负载指标中的6项上达到最低值,在该协议下支持改进的交互安全性。真实Unitree G1舞蹈和站立参考部署定性地展示了有界肢体运动、自主恢复和指令恢复。

英文摘要

Humanoid motion trackers perform reliably within learned tracking distributions, but falls can move the robot into low-height, contact-rich states from which an advancing command is temporarily unreachable. Tracking-only policies may chase infeasible references, producing rapid, large-amplitude limb corrections that increase risk to the robot and its surroundings. We present StableMimic, a unified tracker trained beyond the nominal tracking distribution. Perturbed resets around multiple human get-up references expose prone, supine, off-balance, and intermediate ground-contact states, shaping structured recovery that returns the robot to the trackable region. Because tracking and recovery occupy markedly different state--action distributions, StableMimic uses dedicated experts for each regime and a proprioceptive gate that continuously blends their actions. A hidden successor-state objective teaches human-reference-shaped recovery without exposing reference identity or phase to the deployed Actor; deployment requires no get-up reference, recovery command, trajectory retrieval, or external policy switch. On the complete retargeted LAFAN1 dance subset, StableMimic achieves the lowest errors on all four tracking metrics among five methods. Across 100 matched push-to-fall trials per method, it recovers in 100/100 and attains the lowest values on six of seven post-fall motion and load measures, supporting improved interaction safety under this protocol. Real Unitree G1 dance and standing-reference deployments qualitatively demonstrate bounded limb motion, autonomous recovery, and command resumption.

Comments8 pages, 7 figures. Preprint, not formally peer-reviewed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑