arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向6G的电磁世界模型:环境重建与信道预测的统一框架

Electromagnetic World Model for 6G: A Unified Framework for Joint Environment Reconstruction and Channel Prediction

Yizhu Zhao, Li Yu, Jianhua Zhang, Yuxiang Zhang, Zhen Zhang, Guangyi Liu

arXiv 2608.17769首次发表:更新:

AI 中文总结

该研究针对6G智能终端需同时实现环境感知与通信的需求,提出EMWM统一框架,融合多模态信息完成信道预测与环境重建,性能优于基线方法,具备鲁棒性与零样本泛化能力。

AI 中文摘要

感知、通信与智能的融合正成为第六代(6G)无线系统的关键支撑,智能终端需同时支持高效链路建立与可靠环境感知。然而现有研究多利用感知或通信信息解决单一任务,如信道预测或环境重建。鉴于光信号与射频信号均依赖周围环境,本文提出电磁世界模型(EMWM),首个用于环境重建与信道预测的统一框架,该模型学习通用电磁表示,有望为6G任务提供建模基础。具体而言,将部分信道状态信息(CSI)与多视角红-绿-蓝(RGB)图像编码为CSI令牌与视觉令牌,再由具备局部与全局聚合能力的分层世界模型主干共同处理。基于学习到的表示,采用基于混合专家(MoE)的CSI预测头重建完整CSI,同时深度预测头估计多视角深度图并转换为三维(3D)点云。此外,基于校园数字孪生构建了大规模多模态数据集。实验结果表明,EMWM在CSI预测与环境重建任务上均优于传统神经网络及大语言模型(LLM)基线,CSI预测的平方广义余弦相似度(SGCS)达0.9699,且在不同信噪比(SNR)条件下具备鲁棒性,在28 GHz下具备零样本泛化能力。

英文摘要

The integration of sensing, communication, and intelligence is becoming a key enabler for sixth generation (6G) wireless systems, where intelligent terminals are expected to simultaneously support efficient link establishment and reliable environmental sensing. However, existing studies mainly exploit sensing information or communication information to address a single task, such as channel prediction or environment reconstruction. Motivated by the shared dependence of optical and radio-frequency signals on the surrounding environment, we propose the electromagnetic world model (EMWM), the first unified framework for joint environment reconstruction and channel prediction. EMWM learns a common electromagnetic representation with the potential to provide a modeling foundation for 6G tasks. Specifically, partial channel state information (CSI) and multi-view red-green-blue (RGB) images are encoded into CSI and visual tokens and jointly processed by a hierarchical world-model backbone with local and global aggregation. Based on the learned representation, a mixture-of-experts (MoE)-based CSI prediction head reconstructs the complete CSI, while a depth prediction head estimates multi-view depth maps that are further converted into three-dimensional (3D) point clouds. Moreover, a large-scale multi-modal dataset is constructed based on a campus digital twin. Experimental results show that EMWM outperforms conventional neural network and large language model (LLM) baselines in both CSI prediction and environment reconstruction, achieving a squared generalized cosine similarity (SGCS) of 0.9699 for CSI prediction while demonstrating robustness across different signal-to-noise ratio (SNR) conditions and zero-shot generalization at 28 GHz.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑