arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAMMOTH:一种对缺失模态具有鲁棒性的越野移动多模态端到端策略

MAMMOTH: A Multi-Modal End-to-End Policy for Off-Road Mobility Robust to Missing Modality

Ahaan Kotian, Shivani Subramanyan, Suresh Sundaram

arXiv 2607.12965首次发表:更新:

发表机构

Indian Institute of Science(印度科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对非结构化越野环境自主导航难题,提出MAMMOTH多模态端到端策略。它融合多模态观测,用模态丢弃训练,还采用扩散策略学习联合概率分布,经实验验证性能优越,能提升避撞等能力及对缺失模态的泛化。

AI 中文摘要

在非结构化越野环境中进行可靠的自主导航仍然是一项关键的未解决挑战,因为地形极端多样、光照变化剧烈且传感器严重退化。近期进展将此问题视为可通行性代价地图估计或视觉导航任务,但许多方法严重依赖RGB模态,在不同光照下性能不佳。实现鲁棒泛化需要整合提供补充场景信息的模态,而多模态方法对近乎完美的传感器输入存在刚性依赖。为解决这些限制,我们引入MAMMOTH,一种用于鲁棒越野视觉目标条件导航和无向探索的统一端到端导航策略。具体而言,MAMMOTH有效融合多模态观测(RGB、热成像、3D点云及自我速度),并采用模态丢弃方案训练,使其在推理时能泛化到缺失模态。此外,我们采用扩散策略学习基于物理的轨迹和内在可通行性启发式的联合条件概率分布,MAMMOTH利用此启发式偏好更安全、更平滑的轨迹。我们通过在不同越野环境中的大量真实机器人实验验证了MAMMOTH,包括夜间操作。结果表明其性能优越,在避撞、地形感知规划和对缺失模态的泛化方面有显著改进。这项工作使用的代码和数据集将公开提供。

英文摘要

Reliable autonomous navigation in unstructured off-road environments remains a critical unsolved challenge due to extreme terrain diversity, drastic illumination variations and acute sensor degradation. Recent developments have approached the problem as a traversability costmap estimation or visual navigation task. However, many exhibit heavy reliance on RGB modality, leading to poor performance in varied illumination such as glares, shadows or low ambient light. Achieving robust generalization in such conditions requires integrating modalities that provide supplementary scene information. Such multi-modal methods suffer from a rigid dependency on the presence of near-perfect sensor inputs, leaving them unable to robustly handle sensor degradation or individual modality failure. To address these limitations, we introduce MAMMOTH (MAsking Multi-Modal inputs for Off-road Traversability Heuristic-informed navigation), a unified end-to-end navigation policy for robust off-road visual-goal-conditioned navigation and undirected exploration. Specifically, MAMMOTH efficiently fuses multi-modal observations (RGB, Thermal, 3D Pointcloud and Ego Velocity) and is trained with a modality dropout scheme, enabling it to generalize to missing modalities at inference time. Furthermore, we employ a diffusion policy to learn the joint conditional probability distribution of physically-grounded trajectories and a intrinsic traversability heuristic. MAMMOTH utilizes this heuristic to prefer safer, smoother trajectories. We validate MAMMOTH through extensive real-world robot experiments in distinct off-road environments, including night-time operation. Our results demonstrate superior performance, with significant improvements in collision avoidance, terrain-aware planning and generalization to missing modalities. The code and dataset used for this work will be made publicly available.

CommentsAccepted to IROS 2026 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑