AI 中文总结
DreamTrajectory是面向移动操作的轨迹引导框架,通过联合预测末端执行器轨迹与全身动作块、结合轨迹世界模型的测试时细化,在MS-HAB及真实任务中大幅提升操作成功率。
AI 中文摘要
移动操作要求机器人在不断变化的视角和接触条件下协调底盘与机械臂的运动,其动作空间远大于固定基座操作。现有视觉-语言-动作(VLA)策略存在两方面局限:(i)直接将观测映射到全身动作块,在无显式任务空间运动规划的情况下搜索大动作空间,导致底盘-机械臂协调预测不精确;(ii)开环执行预测的动作块,未检查动作是否能实现策略意图的运动,控制误差与未建模接触会累积成规划运动与实际运动的偏差。本文提出DreamTrajectory,这是一种面向语言条件移动操作的轨迹引导框架,针对上述两个局限各引入一个组件:针对(i),DreamTrajectory在单个动作专家中联合预测意图级末端执行器轨迹与全身动作块,使轨迹明确引导底盘-机械臂动作生成而非隐式存在;针对(ii),轻量轨迹世界模型预测候选动作块会诱导的轨迹,测试时采用“搜索-预测-评分”流程选择与规划轨迹对齐度最高的候选。在MS-HAB上,轨迹引导使平均成功率从32.3%提升至47.5%,测试时细化进一步提升至54.8%,在接触丰富的关节物体任务上提升最大;在三个真实世界移动操作任务上,对应平均成功率分别为63.3%、81.7%和90.0%。
英文摘要
Mobile manipulation requires a robot to coordinate base and arm motion under continuously changing viewpoints and contact conditions, within an action space far larger than that of fixed-base manipulation. Existing Vision-Language-Action (VLA) policies are limited in two respects. (i)They map observations directly to whole-body action chunks, searching this large action space without an explicit task-space motion plan, which makes coordinated base--arm prediction imprecise. (ii)They execute the predicted chunk open-loop, without checking whether the actions can realize the motion the policy intended, so control errors and unmodeled contacts accumulate into a gap between planned and realized motion. We present DreamTrajectory, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation. Addressing(i), DreamTrajectory jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert, so that the trajectory explicitly guides base--arm action generation instead of remaining implicit. Addressing(ii), a lightweight trajectory world model predicts the trajectory that a candidate action chunk would induce, and a test-time search--predict--score procedure selects the candidate best aligned with the planned trajectory. On MS-HAB, trajectory guidance raises average success from 32.3% to 47.5% and test-time refinement further to 54.8%, with the largest gains on contact-rich articulated-object tasks. On three real-world mobile manipulation tasks, the corresponding average success rates are 63.3%, 81.7%, and 90.0%.