GeoHAT: 几何自适应混合动作Transformer用于移动操作
GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation
查看机构详情
- Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出GeoHAT框架,通过轻量级傅里叶空间编码器注入几何信息,并采用混合全身动作解码器分解机械臂与基座动作,在ManiSkill-HAB基准上成功率提升23.7%。
中文摘要 AI 辅助
全身移动操作需要在不断变化的视角下协调移动基座和机械臂,这对几何感知和动作生成提出了挑战。当前的策略要么依赖2D特征,要么依赖缺乏密集空间结构的稀疏3D表示,并且通常将机械臂和基座编码在一个动作向量中,忽略了它们各自不同的控制需求。此外,现有的密集融合策略在噪声深度下可能破坏预训练表示,同时带来沉重的计算开销。我们提出了GeoHAT,一个基于简单原则的端到端扩散框架:几何信息应仅在可靠处注入,且仅在需要处被关注。GeoHAT采用轻量级傅里叶空间编码器,将密集的逐像素3D坐标映射为几何标记,无需额外的3D视觉骨干网络。然后,通过由深度有效性调制的逐标记门控融合,将这些标记选择性地注入视觉基础模型特征中,在保留语义先验的同时丰富空间理解。对于动作生成,混合全身动作解码器将机械臂和基座分解到不同的子空间,并通过稀疏交叉注意力让每个动作模态关注其任务相关的视觉上下文,同时因果时序建模捕获时间步内协调和时间步间依赖。在ManiSkill-HAB仿真基准上的实验表明,GeoHAT实现了79.3%的平均成功率,比最强基线高出23.7%。此外,在多种任务上的真实世界实验也证实了在所有基线上的一致改进。
英文摘要
Whole-body mobile manipulation requires coordinating mobile base and manipulator under shifting viewpoints, posing challenges in geometric perception and action generation. Current policies either rely on 2D features or sparse 3D representations that lack dense spatial structure, and typically encode arm and base within one action vector that ignores their distinct control demands. Moreover, existing dense fusion strategies risk corrupting pretrained representations under noisy depth while incurring heavy computational overhead. We present GeoHAT, an end-to-end diffusion-based framework built on a simple principle: geometry should be injected only where reliable and attended to only where needed. GeoHAT employs a lightweight Fourier spatial encoder that maps dense per-pixel 3D coordinates into geometric tokens without an additional 3D vision backbone. These tokens are then selectively injected into vision foundation model features through per-token gated fusion modulated by depth validity, preserving the semantic prior while enriching spatial understanding. For action generation, a Hybrid Whole-Body Action Decoder decomposes arm and base into distinct subspaces and lets each action modality attend to its task-relevant visual context through sparse cross-attention, while causal temporal modeling captures intra-timestep coordination and inter-timestep dependencies. Experiments on the ManiSkill-HAB simulation benchmark demonstrate that GeoHAT achieves a 79.3% mean success rate, surpassing the strongest baseline by 23.7%. Furthermore, real-world experiments on diverse tasks also confirm consistent improvements over all baselines.