arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于视觉语言模型的图像传输个性化数字语义通信

Personalized Digital Semantic Communication for Image Transmission with Vision-Language Models

Nan Li, Li Zhou, Haijun Wang, Jun Xiong, Haitao Zhao, Jibo Wei

arXiv 2608.14260首次发表:更新:

AI 中文总结

针对现有语义通信方案忽略接收端依赖语义的问题,提出整合VLM与LDM的PDSC框架,构建容量受限的个性化语义率失真问题,实验证实其在带宽受限场景下的源语义一致性与个性化表现优于CDDM、MoS等基线。

AI 中文摘要

语义通信(SC)可实现带宽高效的无线图像传输,但现有多数语义通信方案与用户无关,忽略了依赖接收端的语义信息。为解决该问题,本文提出个性化数字语义通信(PDSC)框架,整合基于视觉语言模型(VLM)的语义编码器与基于潜在扩散模型(LDM)的语义解码器。具体而言,语义编码器从源图像和接收端的历史交互中提取源感知个性化语义令牌,这些令牌经矢量量化为离散语义索引,进一步编码为紧凑的固定长度比特流,以适配数字传输。在接收端,语义解码器基于恢复的语义令牌重建个性化图像。此外,本文构建容量受限的个性化语义率失真问题,并引入语义失真指标,该指标联合表征源语义保真度与用户偏好一致性。实验表明,在带宽受限的无线传输场景下,PDSC相较于CDDM、MoS等最新语义通信基线,在源语义一致性和个性化方面表现更优。

英文摘要

Semantic communication (SC) enables bandwidth-efficient wireless image transmission, but most existing SC schemes are user-agnostic and ignore receiver-dependent semantics. To address this issue, we propose a personalized digital semantic communication (PDSC) framework that integrates a vision-language model (VLM)-based semantic encoder with a latent diffusion model (LDM)-based semantic decoder. Specifically, the semantic encoder extracts source-aware personalized semantic tokens from both the source image and the receiver's historical interactions. These tokens are vector-quantized into discrete semantic indices and further encoded into a compact fixed-length bitstream, enabling compatibility with digital transmission. At the receiver, the semantic decoder reconstructs a personalized image conditioned on the recovered semantic tokens. Furthermore, we formulate a capacity-constrained personalized semantic rate-distortion problem and introduce a semantic distortion metric that jointly characterizes source-semantic fidelity and user-preference alignment. Experiments show that PDSC achieves superior source-semantic consistency and personalization over state-of-the-art SC baselines, including CDDM and MoS, under bandwidth-limited wireless transmission.

CommentsAccepted by IEEE GLOBECOM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑