arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12518eess.SP

用于高精度定位的蜂窝信号构建卷积视觉Transformer

Cellular Signal Constructed Convolutional Vision Transformer for High Accuracy Positioning

Junshi Chen, Xuhong Li, Russ Whiton, Fredrik Tufvesson

首次发表
浏览论文内容

中文总结 AI 辅助

针对复杂蜂窝环境中定位难题,提出混合卷积视觉Transformer(ConViT)架构,整合CNNs局部感受野与Transformer全局注意力,评估多种信号融合策略,经实验验证其性能显著优于基准模型,且有较低参数数量和计算复杂度,还具备可解释性。

中文摘要 AI 辅助

现代蜂窝系统采用宽带宽和大型天线阵列来满足高数据速率需求。通信的高时空分辨率也带来了高精度定位这一附带优势。标准卷积神经网络(CNNs)和视觉Transformer在利用延迟 - 角度域信道表示进行定位方面展现出优异性能,但在低信噪比和严重小区间干扰的复杂蜂窝环境中仍面临实际挑战。本文提出一种混合卷积视觉Transformer(ConViT)架构,它整合了CNNs的局部感受野以抑制局部噪声,并采用Transformer来捕捉不同多径分量间的全局注意力。还评估了多种融合来自多个分布式基站信号的策略。应用带有传感器融合的扩展卡尔曼滤波器进一步减轻模型估计的长尾波动。在城市环境中使用大型天线阵列接收的商业长期演进信号进行了全面验证,该信号存在非视距信号和强小区间干扰。ConViT实现了3.46米的距离均方根误差(RMSE)和2.54度的偏航RMSE,显著优于基准模型,同时保持较低的参数数量和降低的计算复杂度。最后,延迟 - 角度功率分布与Transformer注意力权重之间的对应分析证明了模型的可解释性。

英文摘要

Modern cellular systems employ wide bandwidths and large antenna arrays to meet high data rate requirements. The high spatial and temporal resolution for communication also enables high-accuracy positioning as an ancillary benefit. Standard convolutional neural networks (CNNs) and vision Transformers have demonstrated excellent performance in positioning by leveraging delay-angle domain channel representations. However, they still face practical challenges in complicated cellular environments with low signal-to-noise ratios and severe inter-cell interference. This paper proposes a hybrid convolutional vision Transformer (ConViT) architecture that integrates the local receptive fields of CNNs to suppress local noise and employs Transformers to capture global attention among different multipath components. Various fusion strategies for combining signals from multiple distributed base stations are also evaluated. An extended Kalman filter with sensor fusion is applied to further mitigate long tail fluctuations of model estimates. Comprehensive validation is conducted with commercial long-term-evolution signals received by a large antenna array in urban environments with non line-of-sight signals and strong inter-cell interference. ConViT achieves a distance root mean square error (RMSE) of 3.46 meters and a yaw RMSE of 2.54 degrees, significantly outperforming benchmark models, while maintaining a lower parameter count and reduced computational complexity. Finally, a correspondence analysis between delay-angle power distributions and Transformer attention weights demonstrates the interpretability of the model.

↑