发表机构
Zhejiang University; Wuhan University; Hunan University; D-Robotics(浙江大学; 武汉大学; 湖南大学; D机器人公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对跨具身导航挑战,提出基于模仿学习的具身感知导航框架,通过模块化多阶段设计,预训练构建数据集并引入具身几何,微调设计多模态信息注入机制,有效提高不同具身设置下的导航性能。
AI 中文摘要
跨具身导航是具身智能中的关键挑战。由于具身差异,相同视觉观察对不同智能体可能意味着不同动作,仅靠视觉预测会模糊。现有研究主要依赖强化学习,需大规模交互和精细奖励设计,难以支持可扩展预训练和现实世界适应。基于模仿学习的方法也有限。为此,我们提出基于模仿学习的具身感知导航框架,采用模块化多阶段设计。预训练时,从网络视频构建跨具身导航数据集并引入具身几何作为条件令牌减少动作模糊;微调时,设计基于解耦架构的多模态信息注入机制,包括轨迹增强策略生成高风险样本分别训练空间感知和风险感知校正,明确纳入具身几何用于安全导航。实验结果表明该方法有效提高不同具身设置下的导航性能,证明将具身几何纳入具身导航的有效性。
英文摘要
Cross-embodiment navigation is a key challenge in embodied intelligence. Due to differences in embodiment, the same visual observation may imply different actions for different agents, making prediction ambiguous when relying solely on vision. Existing studies mainly rely on reinforcement learning, which requires large-scale interaction and careful reward design, making it difficult to support scalable pretraining and real-world adaptation. In contrast, imitation-learning-based approaches remain limited. To address these challenges, we propose an imitation-learning-based embodiment-aware navigation framework with a modular multi-stage design. In pretraining, we construct a cross-embodiment navigation dataset from Internet videos and introduce embodiment geometry as conditional tokens to reduce action ambiguity under the same observation. In fine-tuning, we design a multimodal information injection mechanism based on a decoupled architecture. Specifically, we design a trajectory augmentation strategy to generate high-risk samples, which are used to train spatial perception and risk-aware correction separately, thereby explicitly incorporating embodiment geometry for safe navigation. Experimental results show that the proposed method effectively improves navigation performance across different embodiment settings, demonstrating the effectiveness of incorporating embodiment geometry into embodied navigation.