发表机构
Beihang University; Zhongguancun Academy; Northeastern University(北京航空航天大学; 中关村学院; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出面向无人机多模态感知的GAAT模型,引入syncPATC、MG-Sparse-MMA、RA-QCGCL等技术,结合UAVMeta与StateBench数据集,在六个下游任务上实现最优迁移性能,为无人机多模态感知提供支撑。
AI 中文摘要
无人机(UAV)多模态感知整合可见光(RGB)、红外(IR)、合成孔径雷达(SAR)及深度传感器,用于不同条件下的场景理解。然而,光学、分辨率及安装差异常导致实际系统仅能实现全局或图像中心对齐。经分词后,视差、平台运动及镜头畸变会使跨模态对应块中心发生偏移,削弱密集对比学习与跨模态融合所依赖的空间对应关系。本文提出GAAT(Geometry-Aware Alignment Transformer,几何感知对齐Transformer),这是一种以对齐为核心的预训练模型,在跨模态交互前先估计局部对应可靠性。GAAT引入syncPATC,无需对应标注即可学习同步视图变换下的块中心一致性;输出几何先验,包括 token 与查询置信度、查询中心及子 token 偏移,以识别残留失配下的可靠局部锚点。在这些先验引导下,MG-Sparse-MMA 基于前 K_s 个可靠区域执行查询介导的稀疏融合,用几何校准的局部更新替代全块密集交互。RA-QCGCL 通过可靠块到块、块到查询及查询到查询的对比分支,将预训练监督与该稀疏查询瓶颈对齐。本文还引入UAVMeta与StateBench,二者提供源自平台遥测及图像统计的四个采集状态评分:相机可靠性、观测尺度、视点稳定性及飞行动作复杂度。针对六个下游任务的大量实验表明,GAAT 具备始终优异的迁移性能,确立其为无人机多模态感知领域的当前最优基础模型;StateBench 还可实现对真实世界采集条件的系统性诊断。
英文摘要
Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in optics, resolution, and mounting often limit practical systems to global or image-center alignment. After tokenization, parallax, platform motion, and lens distortion can shift corresponding patch centers across modalities, weakening the spatial correspondence assumed by dense contrastive learning and cross-modal fusion. We propose GAAT (Geometry-Aware Alignment Transformer), an alignment-first pretrained model that estimates local correspondence reliability before cross-modal interaction. GAAT introduces syncPATC, which learns patch-center consistency under synchronized view transformations without correspondence annotations. It emits geometric priors, including token and query confidence, query centers, and sub-token offsets, that identify reliable local anchors across residual misalignment. Guided by these priors, MG-Sparse-MMA performs query-mediated sparse fusion over top-K_s reliable regions, replacing dense all-patch interaction with geometry-calibrated local updates. RA-QCGCL aligns pretraining supervision with this sparse query bottleneck through reliable patch-to-patch, patch-to-query, and query-to-query contrastive branches. We introduce UAVMeta and StateBench, which provide four acquisition-state scores derived from platform telemetry and image statistics: camera reliability, observation scale, viewpoint stability, and flight maneuver complexity. Extensive experiments across six downstream tasks demonstrate consistently superior transfer performance, establishing GAAT as a state-of-the-art multimodal foundation model for UAV perception. StateBench further enables a systematic diagnosis of real-world acquisition conditions.