AI 中文总结
本研究提出RoboReact框架,通过生成自我中心视频并结合视觉语言引导的闭环控制,实现人形机器人全身操控技能的通用获取,可在多样场景中稳健执行。
AI 中文摘要
人形机器人具备在人类环境中执行灵巧操控的潜力,但由于硬件数据采集成本高昂、标注劳动密集,获取多样且通用的技能仍代价不菲。视频生成模型的最新进展为从视觉观测中合成丰富的操控经验提供了契机,不过将这类想象的行为转化为可执行的全身人形机器人技能,在很大程度上仍未被探索。本研究提出RoboReact框架,可从单一自我中心RGB-D观测中自动合成全身人形机器人操控技能。RoboReact生成人类操控视频,通过感知深度的3D重建提取保留几何特征的交互关键帧,并将其重定向至高自由度人形机器人平台,同时保留手-物交互几何。为弥合想象规划与物理执行之间的差距,RoboReact执行以物体为中心的在线重定位,并利用视觉语言模型引导的优化循环,在几何不匹配和执行偏差下适配技能。优化后的技能通过全身控制器执行,实现协调的全身操控与灵巧交互。在真实人形机器人上的实验表明,RoboReact可在多样物体配置间泛化,且无需遥操作或人类演示即可从执行干扰中稳健恢复。这些结果凸显了将生成模型、视觉语言推理与闭环控制相结合,用于可扩展人形机器人技能获取的潜力。
英文摘要
Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive annotation. Recent advances in video generative models provide a promising opportunity to synthesize rich manipulation experiences from visual observations, but transferring such imagined behaviors into executable whole-body humanoid skills remains largely unexplored. In this work, we present RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation. RoboReact generates human manipulation videos, extracts geometry-preserving interaction keyframes through depth-aware 3D reconstruction, and retargets them to high-DoF humanoid platforms while preserving hand-object interaction geometry. To bridge the gap between imagined plans and physical execution, RoboReact performs online object-centric re-grounding and leverages a vision-language model-guided refinement loop to adapt skills under geometric mismatch and execution deviations. The refined skills are executed through a whole-body controller, enabling coordinated whole-body manipulation and dexterous interaction. Experiments on real humanoid robots demonstrate that RoboReact generalizes across diverse object configurations and robustly recovers from execution disturbances without requiring teleoperation or human demonstrations. These results highlight the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.