MoPA:通过子系统特定感知对齐实现协调移动操作
MoPA: Coordinated Mobile Manipulation via Subsystem-Specific Perception Alignment
浏览论文内容
中文总结 AI 辅助
MoPA通过双感知流和感知到动作适应,对齐移动与操作的感知条件化,在ManiSkill-HAB和真实任务中实现最先进性能。
中文摘要 AI 辅助
移动操作需要针对底盘运动和手臂控制在不同空间尺度上的感知证据,而这两种动作模态在运动学上保持耦合。现有策略通常对不同子系统采用专门的动作生成,但将异构动作分支条件化于共享的感知表示上,使得子系统特定的感知-动作对应关系隐含不清。我们提出MoPA,一个在动作层面保持协调的同时,将感知条件化与移动和操作对齐的框架。双感知流采用两个相互掩码的查询库,从共享的视觉-语言上下文中提取分离的感知表示。感知到动作适应在结构化混合Transformer解码器的每一层联合更新每个查询库及其对应的动作流,同时实现两个动作流之间的信息交换。耦合条件流匹配学习一个联合向量场,以协调生成两个动作块。在ManiSkill-HAB基准上,MoPA在所有三个任务套件中均取得了最先进的性能。在四个真实世界任务中,MoPA实现了76.3%的平均完整任务成功率,比最佳基线高出12.5个百分点。消融研究和进一步分析验证了所提出设计的有效性。网站可访问:此https URL。
英文摘要
Mobile manipulation requires perceptual evidence at different spatial scales for base motion and arm control, while the two action modalities remain kinematically coupled. Existing policies often employ specialized action generation for different subsystems but condition heterogeneous action branches on a shared perceptual representation, leaving subsystem-specific perception-action correspondence implicit. We present MoPA, a framework that aligns perceptual conditioning with mobility and manipulation while preserving coordination at the action level. Dual Perceptual Streams employ two mutually masked query banks to extract separate perceptual representations from a shared vision-language context. Perception2Action Adaptation jointly updates each query bank and its corresponding action stream at every layer of a structured Mixture-of-Transformers decoder, while enabling information exchange between the two action streams. Coupled conditional flow matching learns a joint vector field for coordinated generation of both action chunks. On the ManiSkill-HAB benchmark, MoPA achieves state-of-the-art performance across all three task suites. Across four real-world tasks, MoPA achieves a mean full-task success rate of 76.3%, outperforming the best baseline by 12.5 percentage points. Ablation studies and further analyses validate the effectiveness of the proposed design. Website is available at: https://mopa-policy.github.io/.
发表机构
- THU(清华大学)
- Xspark AI
- HKUST (GZ)(香港科技大学(广州))
- HKU(香港大学)
机构由 AI 辅助整理,请以论文原文为准。