AI 中文总结
针对车联网中高维多模态传感数据传输面临的频谱瓶颈,提出AirTF框架。利用视觉变压器编码器提取语义令牌,通过共享无线信道并发传输,利用多址信道叠加特性空中融合多模态语义,提升频谱效率。
AI 中文摘要
在车联网(IoV)中,将高维多模态传感数据传输到边缘服务器以完成对时间敏感的任务面临着严重的频谱瓶颈。为了解决这个问题,我们提出了一种基础模型驱动的空中令牌融合(AirTF)框架,用于面向任务的多模态令牌通信。与现有的依赖具有有限局部感受野的卷积神经网络(CNN)的分割方案不同,AirTF利用视觉变压器(ViT)编码器从分布式异构传感器中提取全局上下文语义令牌。通过在共享无线信道上并发传输这些空间对齐的令牌,我们的框架利用多址信道的叠加特性直接在空中固有地融合互补的多模态语义(例如,RGB和红外)。与正交传输相比,这种机制显著提高了频谱效率。此外,预训练基础模型的集成提供了关键的视觉先验,有效地解决了ViT在有限的特定场景语义分割数据集上对数据的高需求。实验表明,在AWGN和衰落信道上,AirTF始终优于正交传输和基于CNN的融合基线。在三用户设置、残余同步误差和不完美信道状态信息估计下的额外评估进一步证实了其鲁棒性。源代码将在接受后公开提供
英文摘要
In the Internet of Vehicles (IoV), transmitting high-dimensional multi-modal sensory data to edge servers for time-sensitive tasks faces severe spectrum bottlenecks. To address this, we propose a foundation model-driven over-the-air token fusion (AirTF) framework for task-oriented multi-modal token communications. Unlike existing schemes for segmentation that rely on convolutional neural networks (CNNs) with limited local receptive fields, AirTF leverages vision transformer (ViT) encoders to extract globally contextualized semantic tokens from distributed heterogeneous sensors. By concurrently transmitting these spatially aligned tokens over a shared wireless channel, our framework exploits the superposition property of the multiple access channel to inherently fuse complementary multi-modal semantics (e.g., RGB and infrared) directly over the air. This mechanism significantly enhances spectral efficiency compared to orthogonal transmission. Furthermore, the integration of a pre-trained foundation model provides critical visual priors, effectively addressing the data-hungry nature of ViTs on limited, scenario-specific semantic segmentation datasets. Experiments demonstrate that AirTF consistently outperforms orthogonal transmission and CNN-based fusion baselines across AWGN and fading channels. Additional evaluations under a three-user setting, residual synchronization errors, and imperfect channel state information estimation further confirm its robustness. The source code will be made publicly available upon acceptance.
CommentsManuscript under review