arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PredActor:用于可操控机载人形机器人控制的预测性动作扩散

PredActor: Predictive Action Diffusion for Steerable Onboard Humanoid Control

Lei Ye, Haibo Gao, Yitang Li, Peng Xu, Zetong Jing, Junhan Sun, Fanrong Dong, Ziqi Han, Xue Wang, Jianhua Sun, Cewu Lu, Hao Zhao, Liang Ding

arXiv 2609.24840首次发表:更新:

发表机构

Harbin Institute of Technology; Shanghai Innovation Institute; RoboParty Lab; Tsinghua University; Shanghai Jiao Tong University(哈尔滨工业大学; 上海创新研究院; RoboParty实验室; 清华大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PredActor提出预测性动作扩散策略,联合生成动作与内部未来状态轨迹,通过分类器引导实现测试时运动操控,在仿真和Unitree G1硬件上验证了文本条件运动、干扰响应与语义插值能力。

AI 中文摘要

扩散模型提供了灵活的运动生成能力,但将这种灵活性转化为具有反馈响应的人形机器人控制仍然具有挑战性。分层系统通过可能超出独立跟踪器能力的参考来引导运动,将恢复和物理执行主要留给跟踪器。仅动作扩散直接生成动作,但缺乏用于测试时运动目标的显式未来状态轨迹。联合状态-动作扩散提供了这种表示,然而具有代表性的控制器通常依赖于特权全身状态,并且对学习行为选择和测试时运动操控的支持仍然分散。我们提出了PredActor,一种预测性动作扩散策略,通过本体感觉观测将互补的操控能力整合到一个直接执行的策略中。以本体感觉历史和可选任务上下文为条件,PredActor联合生成可执行动作和内部未来状态轨迹。无分类器引导增强了文本条件行为,而分类器引导则将预测状态导向测试时目标。仅执行动作,无需单独的运动参考跟踪器或外部估计的全身状态作为策略输入。在仿真中,PredActor到达所有15个目标目的地,并实现0.580的文本检索分数,而条件动作扩散为0.373,且观测到的干扰生存率相似。为了使这种引导策略在机载上实用,滚动去噪和计算保留运行时优化将完整回调减少到Jetson Orin NX上的中位数16.790毫秒和p95 19.383毫秒,均低于20毫秒控制周期。我们在Unitree G1上部署了PredActor;跨仿真和物理硬件的评估展示了文本条件运动、干扰响应、操纵杆控制和语义插值。

英文摘要

Diffusion models provide a flexible framework for motion generation, but turning this flexibility into closed-loop humanoid control remains challenging. Hierarchical generator-tracker systems steer motion through reference trajectories, yet these references may exceed the capabilities of the downstream tracker, leaving physical feasibility and disturbance recovery largely to a separate control module. Action-only diffusion avoids this separation by directly generating executable actions, but provides no explicit future-state trajectory that can be steered toward test-time motion objectives. Joint state-action diffusion offers a natural alternative, but existing controllers often rely on privileged full-body states, while learned behavior selection and test-time motion steering remain only partially integrated. We present PredActor, a predictive action diffusion policy that unifies both steering modes in one directly executed policy using proprioception alone. Given proprioceptive history and optional task context, PredActor jointly predicts actions and an internal future-state trajectory that enables guidance: classifier-free guidance strengthens text-conditioned motion, while classifier guidance steers future states toward test-time objectives. Only actions are executed, requiring neither a motion-reference tracker nor privileged full-body states. In simulation, PredActor reaches 44 of 45 destination targets and achieves a text retrieval score of 0.539 versus 0.424 for conditional action diffusion, with similar disturbance survival. Rolling denoising and computation-preserving runtime optimizations reduce the complete callback to 16.790 ms median and 19.383 ms p95 on a Jetson Orin NX, within the 20 ms control period. Deployed on a Unitree G1, PredActor demonstrates text-conditioned motion, disturbance response, joystick control, and semantic interpolation in simulation and hardware.

CommentsProject page: https://masteryip.github.io/predactor.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑