M3GA-Wild:用于森林中多模态、多会话地空地点识别的大规模数据集与基准
M3GA-Wild: A Large-Scale Dataset and Benchmark for Multi-Modal Multi-session Ground-to-Aerial Place Recognition in Forests
浏览论文内容
中文总结 AI 辅助
M3GA-Wild是首个森林多模态多会话地空地点识别基准,含36公里地面和370公顷航空数据,揭示LiDAR优于视觉及跨模态对齐挑战。
中文摘要 AI 辅助
我们提出了M3GA-Wild,这是森林中多模态、多会话地空地点识别的首个基准。M3GA-Wild统一并扩展了现有的森林定位数据集,提供了一个全面的基准,包含来自地面遍历的同步RGB图像和LiDAR数据,覆盖36公里,对齐的高分辨率航空图像和多高度LiDAR数据,覆盖370公顷,以及精确的地理参考6自由度位姿,用于精确评估。M3GA-Wild捕捉了具有不同视角、遮挡和环境条件的多样化森林场景,使得能够系统评估视觉、LiDAR、跨模态和多模态方法。基线实验表明,在严重的视角差异下,基于LiDAR的方法显著优于仅视觉的方法,而当前的多模态融合策略由于跨模态对齐不佳而收益有限。通过将航空RGB图像与地理参考的航空LiDAR配对,M3GA-Wild还使得能够评估用于单目深度估计的基础模型,作为从森林图像获取3D几何的廉价来源,初步实验揭示了当前方法的不足。这些结果突出了跨平台定位中的关键挑战,包括模态不对齐和严重的领域差距。M3GA-Wild建立了一个新的基准,以支持在非结构化自然环境中鲁棒多模态定位和长期自主性的研究。数据集和代码将在录用后提供。
英文摘要
We present M3GA-Wild, the first benchmark for multi-modal, multi-session ground-to-aerial place recognition in forests. M3GA-Wild unifies and extends existing forest localisation datasets, providing a holistic benchmark with synchronised RGB imagery and LiDAR from ground traversals spanning 36 km, aligned high-resolution aerial imagery and multi-altitude LiDAR covering 370 hectares, and accurate geo-referenced 6-DoF poses for precise evaluation. M3GA-Wild captures diverse forest scenes with varying viewpoints, occlusion, and environmental conditions, enabling systematic evaluation of visual, LiDAR, cross-modal, and multi-modal methods. Baseline experiments show that LiDAR-based approaches significantly outperform vision-only methods under severe viewpoint differences, while current multi-modal fusion strategies yield limited gains due to poor cross-modal alignment. By pairing aerial RGB imagery with geo-referenced aerial LiDAR, M3GA-Wild also enables evaluation of foundation models for monocular depth estimation as a cheap source of 3D geometry from forest imagery, with initial experiments revealing shortfalls of current methods. These results highlight key challenges in cross-platform localisation, including modality misalignment and severe domain gaps. M3GA-Wild establishes a new benchmark to support research in robust multi-modal localisation and long-term autonomy in unstructured natural environments. The dataset and code will be available upon acceptance.
发表机构
- Queensland University of Technology (QUT)(昆士兰科技大学)
- CSIRO Robotics(澳大利亚联邦科学与工业研究组织机器人部门)
机构由 AI 辅助整理,请以论文原文为准。