发表机构
Great Bay University; Great Bay Institute for Advanced Study (GBIAS); Guangdong University of Technology; Chalmers University of Technology; Khalifa University(大湾区大学; 大湾区高等研究院; 广东工业大学; 查尔姆斯理工大学; 哈利法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对复杂城市环境中的多用户定位问题,提出一种基于跨模态Transformer的视觉-无线融合框架,利用导频索引CSI与视觉记忆,通过交叉注意力实现用户级定位,实验验证其优于多种基线方法。
AI 中文摘要
在复杂的城市环境中,准确的多用户定位具有挑战性,无线测量在噪声、阻塞和多径条件下可能变得模糊,而视觉观测提供了互补的空间上下文。本文提出了一种利用导频索引信道状态信息(CSI)的视觉-无线融合框架,用于多用户定位。正交导频索引在CSI令牌序列和定位输出中保留了通信UE的身份。该模型将每个导频索引的CSI观测编码为查询令牌,并使用交叉注意力从空间视觉记忆中检索用户特定信息。CSI令牌之间的自注意力进一步捕捉用户间交互,而由此产生的多模态表示用于逐用户定位。在不同数据集上的实验表明,该模型相对于基于模型、仅CSI和多模态融合基线方法均取得了一致的改进。进一步的实验评估了该模型在不同无线和视觉条件下的性能。
英文摘要
Accurate multi-user localization is challenging in complex urban environments, where wireless measurements can become ambiguous under noise, blockage, and multipath, while visual observations provide complementary spatial context. This paper presents a vision-wireless fusion framework for multi-user localization using pilot-indexed channel state information (CSI). Orthogonal pilot indices preserve the identities of the communicating UEs in the CSI token sequence and localization outputs. The model encodes each pilot-indexed CSI observation as a query token and uses cross-attention to retrieve user-specific information from spatial visual memory. Self-attention among CSI tokens further captures inter-user interactions, while the resulting multimodal representations are used for user-wise localization. Experiments on different datasets show consistent improvements over model-based, CSI-only, and multimodal-fusion baselines. Further experiments evaluate the model under different wireless and visual conditions.
Comments12 pages, 8 figures, 7 tables