arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

M3F-UAV:一种用于低空无线传感的缺失模态多模态基础模型

M3F-UAV: A Missing-Modality Multimodal Foundation Model for Low-Altitude Wireless Sensing

Pengxuan Gao, Kai Ying, Botao Wu, Jianhua Mo, Qingsong Wen

arXiv 2607.13678首次发表:更新:

AI 中文总结

针对复杂环境下单模态模型可靠性低的问题,提出M3F-UAV模型,通过特定模态预训练特征提取器、跨模态融合及缺失模态感知预训练,从多观测中学习统一表示,在LAMBDA数据集实验中性能优于单模态基线且在缺失模态下稳健。

AI 中文摘要

低空无人机正成为无线智能任务的关键平台。但实际的低空无线系统通常在复杂城市环境中运行,视觉遮挡、稀疏几何观测、多径传播和传感器故障会降低单模态模型的可靠性。本文提出M3F-UAV,一种用于低空无线传感的缺失模态多模态基础模型。该框架从视觉、几何和无线观测中学习统一的多模态表示。具体采用特定模态的预训练特征提取器分别处理RGB/深度图像、激光雷达点云和CSI矩阵。通过跨模态融合和缺失模态感知预训练,M3F-UAV能从不同模态组合中提取固定大小特征并通过轻量级任务头适应下游低空无线任务。在LAMBDA数据集上的实验表明,M3F-UAV优于单模态基线且在缺失模态设置下保持稳健性能。

英文摘要

Low-altitude unmanned aerial vehicles (UAVs) are emerging as key platforms for wireless intelligence tasks. However, practical low-altitude wireless systems usually operate in complex urban environments, where visual occlusion, sparse geometric observations, multipath propagation, and sensor failures may degrade the reliability of single-modality models. To address these challenges, this paper proposes M3F-UAV, a missing-modality multimodal foundation model for low-altitude wireless sensing. The proposed framework learns a unified multimodal representation from visual, geometric, and wireless observations. Specifically, modality-specific pretrained feature extractors are adopted for RGB/depth images, LiDAR point clouds, and CSI matrices, respectively. Through cross-modal fusion and missing-modality-aware pretraining with feature-level masked reconstruction and UAV localization objectives, M3F-UAV can extract fixed-size features from different modality combinations and adapt them to downstream low-altitude wireless tasks with lightweight task heads. Experiments on the LAMBDA dataset show that M3F-UAV outperforms single-modality baselines and maintains robust performance under missing-modality settings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑