发表机构
The Edward S. Rogers Sr. Department of Electrical and Computer Engineering, University of Toronto(多伦多大学爱德华·S·罗杰斯Sr.电气与计算机工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种整合相机、LiDAR和RF导频的多模态学习框架,用于混合LoS/NLoS环境下的MU-MISO波束成形优化,通过图像引导LiDAR特征与RF联合处理,实现比基准更高的最小速率。
AI 中文摘要
本文提出了一种多模态学习框架,该框架整合了相机图像、LiDAR点云和射频(RF)导频,以在混合视距(LoS)和非视距(NLoS)环境中实现多用户多输入单输出(MU-MISO)波束成形优化。在LoS主导的环境中,可以利用与用户位置相关的传感信息来设计波束成形。然而,在现实动态环境中,由于遮挡,无法保证所有用户都存在LoS路径。为了表征NLoS传播,需要周围障碍物的详细三维(3D)信息。本文整合相机图像和LiDAR点与RF导频,以捕获用户和周围物体的详细3D几何、物体类别和材料信息。所获得的多模态信息为直接LoS路径和反射NLoS传播提供了有用的线索,并为波束成形设计提供了丰富的信道相关信息。具体而言,首先使用边界框检测神经网络从相机图像中提取二维空间、物体类别和材料信息。然后,图像导出的信息指导LiDAR点的选择,以获取用户和周围障碍物的3D几何信息,并为选定的点分配物体类别和材料标签。图像引导的LiDAR特征与RF导频一起,由设计的神经架构联合处理,以优化波束成形向量。仿真结果表明,所提出的方法比基准方法实现了更高的最小速率。
英文摘要
This paper proposes a multimodal learning framework that integrates camera images, LiDAR point clouds, and radio-frequency (RF) pilots to realize multi-user multiple-input single-output (MU-MISO) beamforming optimization in mixed line-of-sight (LoS) and non-line-of-sight (NLoS) environments. In LoS-dominant environments, beamforming can be designed by utilizing sensing information related to user locations. However, LoS paths cannot be guaranteed to exist for all users in realistic dynamic environments due to blockages. To characterize NLoS propagation, detailed three-dimensional (3D) information of surrounding obstacles is required. This paper integrates camera images and LiDAR points with RF pilots to capture detailed 3D geometric, object-class, and material information of users and surrounding objects. The obtained multimodal information provides useful cues for both direct LoS paths and reflected NLoS propagation, and provides rich channel-related information for beamforming design. Specifically, a bounding-box detection neural network is first used to extract two-dimensional spatial, object-class, and material information from camera images. The image-derived information then guides the selection of LiDAR points to obtain 3D geometric information of users and surrounding obstacles, and assigns object-class and material labels to the selected points. The image-guided LiDAR features, together with RF pilots, are jointly processed by a designed neural architecture to optimize the beamforming vectors. Simulation results show that the proposed method achieves a higher minimum rate than benchmarks.