arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

采样、仿真、筛选:无需训练的物理在环人形机器人文本到运动生成

Sample, Simulate, Select: Physics-in-the-Loop Text-to-Motion for Humanoids Without Training

Raphael Memmesheimer, Sven Behnke

arXiv 2609.26420首次发表:更新:

发表机构

University of Bonn; Lamarr Institute for Machine Learning and Artificial Intelligence(波恩大学; 拉马尔机器学习和人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出S$^3$方法,将冻结的文本到运动模型与仿真验证结合,无需训练即可提升人形机器人运动生成的可执行性,在HumanML3D上显著提高直立执行率,并验证了硬件可行性。

AI 中文摘要

文本到运动模型能生成看似合理的人体运动,但未对机器人的动力学进行建模;全身跟踪控制器能可靠地执行机器人参考轨迹,但无法重新规划不可行的参考轨迹。近期的语言到人形机器人系统通过训练来弥合这一差距。我们通过将部署控制器本身置于循环中,衡量在不进行任何训练的情况下能弥合多大差距。采样-仿真-筛选(S$^3$)从冻结的文本到运动模型中为每个提示词抽取$N$个运动,通过方向匹配逆运动学将每个运动重定向到Unitree G1,在完整刚体动力学下使用预训练的SONIC跟踪策略对所有运动进行滚动仿真,并保留策略执行效果最佳的候选。由于验证器是确定性仿真器本身,S$^3$通过构造达到了任意$N$上限;我们衡量的是该上限位于何处以及哪些因素未达到该上限。在200个分层HumanML3D测试提示词上,$N=8$时,直立执行率从83.5%提升至89.5%,硬件门控通过数从33提升至85;在完整测试集(4,184个提示词)上,直立执行率从80.5%提升至89.5%。一个能很好预测跌倒的动力学验证器(AUROC 0.90)仅恢复了这一增益的四分之一:对提示词自身的候选进行排序比分类总体更难。选择无法修复的一类问题是降低骨盆的提示词,而一个在重定向机器人数据上训练的生成器能够执行这类提示词。我们进一步使用标准文本-运动评估器对执行运动的语义保真度进行评分,并设置真实动作捕捉对照组以将损失归因于机器人投影,将重定向器与GMR进行消融比较(互补失败:两者的任意8上限提升至95.0%),并在真实G1上执行所有177个门控选择的片段:每个片段都以站立完成,硬件跟踪误差与仿真匹配($r=0.94$)。

英文摘要

Text-to-motion models generate plausible human motion but do not model a robot's dynamics; whole-body tracking controllers execute robot references reliably but cannot replan an infeasible one. Recent language-to-humanoid systems bridge this gap by training. We measure how much of the gap closes with no training at all, by putting the deployment controller itself in the loop. Sample-simulate-select (S$^3$) draws $N$ motions per prompt from a frozen text-to-motion model, retargets each to a Unitree G1 by direction-matching inverse kinematics, rolls all of them out under full rigid-body dynamics with the pretrained SONIC tracking policy, and keeps the candidate the policy executed best. Because the verifier is the deterministic simulator itself, S$^3$ attains the any-of-$N$ ceiling by construction; what we measure is where that ceiling lies and what falls short of it. On 200 stratified HumanML3D test prompts with $N=8$, upright execution rises from 83.5% to 89.5% and hardware-gate passes from 33 to 85; on the complete test split (4,184 prompts) it rises from 80.5% to 89.5%. A kinematic verifier that predicts falls well (AUROC 0.90) recovers only a quarter of this gain: ranking a prompt's own candidates is harder than classifying the population. What selection cannot fix is one class, prompts that lower the pelvis, which a generator trained on retargeted robot data does execute. We further score the semantic fidelity of the executed motion with the standard text-motion evaluator, with a real-mocap control that attributes the loss to the robot projection, ablate the retargeter against GMR (complementary failures: the any-of-8 ceiling rises to 95.0% over both), and execute all 177 gate-selected clips on the real G1: every one completes standing, with hardware tracking error matching simulation ($r=0.94$).

Comments8 pages, 9 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑