发表机构
Univ Rennes, CNRS, IETR, UMR 6164(雷恩大学,法国国家科学研究中心,信息与电子工程研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MLLM多用户网络中OFDM传输的PAPR问题,提出MAPS方案,通过两阶段训练联合优化跨模态对齐与PAPR降低,在3dB IBO下AVQA准确率提升64.5%,能量效率提升64.4%。
AI 中文摘要
由多模态大语言模型(MLLM)赋能的面向任务的语义通信(SemCom)最近作为一种高效传输多模态信息的有前景范式而出现,然而通过OFDM信道传输令牌嵌入面临由硬件损伤带来的关键挑战,特别是高峰均功率比(PAPR)在非线性功率放大器下严重降低能量效率。在本文中,我们提出了一种新颖的PAPR感知的面向任务的多模态令牌传输框架。具体而言,我们提出了MAPS(多模态AI驱动的PAPR感知令牌传输)方案,用于能量高效的多用户无线通信。关键挑战是联合确保模态间一致性、任务相关令牌传输和PAPR降低。为解决这一问题,我们采用两阶段训练策略,该策略整合了联合跨模态对齐和PAPR降低,随后进行面向任务的微调。MAPS在现实的Modified Rapp功率放大器下通过OFDM信道传输多模态令牌嵌入。具体来说,我们在文本分支中引入可训练的线性投影以实现有效的梯度传播,设计具有基于方差的用户间均衡的平衡多模态PAPR损失,并将重建损失锚定到固定语义目标以保持跨模态对齐。仿真结果表明,MAPS在所有模态上实现了均衡的PAPR降低,同时保持强模态间一致性。在功率放大器非线性下,在输入回退(IBO)为3 dB时,MAPS在音视频问答(AVQA)准确率(64.5%增益)和面向任务的能量效率(64.4%增益)方面均优于基线方案,突显了其在能量高效多模态语义通信中的有效性。
英文摘要
Task-oriented semantic communication (SemCom) empowered by multimodal large language models (MLLMs) has recently emerged as a promising paradigm for efficiently transmitting multimodal information, yet transmitting token embeddings over OFDM channels faces critical challenges due to hardware impairments, particularly the high peak-to-average power ratio (PAPR) that severely degrades energy efficiency under nonlinear power amplifiers. In this paper, we propose a novel PAPR-aware task-oriented multimodal token transmission framework. Specifically, we propose MAPS (Multimodal AI-driven PAPR-aware Token Transmission) scheme for energy-efficient multiuser wireless communications. The key challenge is to jointly ensure inter-modal consistency, task-relevant token transmission, and PAPR reduction. To address this, we adopt a two-stage training strategy that integrates joint cross-modal alignment and PAPR reduction, followed by task-oriented fine-tuning. MAPS transmits multimodal token embeddings over an OFDM channel under a realistic Modified Rapp power amplifier. Specifically, we incorporate a trainable linear projection in the text branch for effective gradient propagation, design a balanced multimodal PAPR loss with variance-based equalization across users, and anchor the reconstruction loss to a fixed semantic target to preserve cross-modal alignment. Simulation results show that MAPS achieves balanced PAPR reduction across all modalities while maintaining strong inter-modal consistency. Under power amplifier nonlinearity, MAPS outperforms baseline schemes in both audio-visual question answering (AVQA) accuracy (64.5% gain) and task-oriented energy efficiency (64.4% gain) at an input back-off (IBO) of 3 dB, highlighting its effectiveness for energy-efficient multimodal semantic communication.
Comments5 pages, 9 figures. Accepted at IEEE PIMRC 2026