arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

由大模型和语义感知混合波束成形赋能的面向查询的通用图像语义编码

Generalized Query-Oriented Image Semantic Coding Empowered by Large AI Models and Semantic-Aware Hybrid Beamforming

Sin-Yu Huang, Vincent W. S. Wong

arXiv 2607.28276首次发表:更新:

AI 中文总结

该研究针对现有语义编码未考虑用户意图、泛化能力受限等问题,提出由大模型与SA-HBF算法赋能的QO-ISC框架,其在未见过的对象类别上的性能优于传统编解码器及两种先进方案。

AI 中文摘要

语义通信是一种新兴范式,可在传输过程中保留数据的语义含义。然而,人类用户往往会根据自身意图关注特定语义内容,而当前的语义编码设计通常未考虑用户意图。此外,大多数现有语义模型是使用特定数据集进行微调的,这限制了它们的泛化能力。再者,如何在大规模多输入多输出正交频分复用(MIMO-OFDM)系统中对语义重要的特征进行优先级排序,这一问题在很大程度上尚未得到探索。为应对上述挑战,本文提出一种面向查询的通用图像语义编码(QO-ISC)框架。在该框架中,发射机提取与用户查询相关的特征,接收机基于这些特征重建图像。我们使用预训练的大人工智能模型(LAM)来增强通用特征表示,并开发了语义感知混合波束成形(SA-HBF)算法,以便为大规模MIMO-OFDM系统对语义重要的特征进行优先级排序。在数据集内未见过的对象类别上进行评估时,仿真结果表明,我们提出的通用QO-ISC框架的性能优于传统编解码器和两种最先进的语义编码方案。

英文摘要

Semantic communication is an emerging paradigm that can preserve the meaning of data during transmission. However, human users are often interested in specific semantic content based on their intent, and users' intent is often not considered in current semantic coding design. Moreover, most of the existing semantic models are fine-tuned using specific datasets, which limits their generalization capability. Furthermore, how to prioritize semantically important features in large-scale multiple-input multiple-output orthogonal frequency-division multiplexing (MIMO-OFDM) systems remains largely unexplored. To address the aforementioned challenges, in this paper, we propose a generalized query-oriented image semantic coding (QO-ISC) framework. In the proposed framework, the transmitter extracts features which are relevant to the user's query and the receiver reconstructs an image based on those features. We use a pretrained large artificial intelligence (AI) model (LAM) to enhance general feature representations. We develop a semantic-aware hybrid beamforming (SA-HBF) algorithm to prioritize semantically important features for large-scale MIMO-OFDM system. When evaluated on unseen object categories within the dataset, simulation results show that our proposed generalized QO-ISC framework achieves better performance than the traditional codec and two state-of-the-art semantic coding schemes.

CommentsAccepted by IEEE Transactions on Communications (TCOM)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑