arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.12808cs.ITcs.CVcs.LGmath.IT

联合源-信道-生成编码:从以失真为导向的重建到语义一致的生成

Joint Source-Channel-Generation Coding: From Distortion-oriented Reconstruction to Semantic-consistent Generation

  • Cooperative Medianet Innovation Center, Shanghai Jiao Tong University, Shanghai, China(合作中位网创新中心,上海交通大学,上海,中国)

机构由 AI 辅助整理,请以论文原文为准。

Tong Wu, Zhiyong Chen, Guo Lu, Li Song, Feng Yang, Meixia Tao, Wenjun Zhang

更新

AI总结:

本文提出JSCGC,通过概率生成替代确定性重建,提升图像传输的感知质量和语义保真度,优于传统JSCC方法。

AI中文摘要:

传统通信系统,包括基于分离的编码和由AI驱动的联合源-信道编码(JSCC),主要受到香农的率失真理论的指导。然而,依赖通用失真度量标准无法捕捉复杂的视觉感知,常常导致模糊或不真实的重建。在本文中,我们提出联合源-信道-生成编码(JSCGC),一种新的范式,将重点从确定性重建转向概率生成。JSCGC在接收端利用生成模型作为生成器而非传统解码器来参数化数据分布,使在信道约束下直接最大化互信息的同时,通过控制随机采样生成高保真的输出,这些输出位于真实数据流形上。我们进一步推导了在给定传输互信息下的最大语义不一致的理论下限,阐明了在控制生成过程方面的通信基本限制。对图像传输的大量实验表明,JSCGC显著提高了感知质量和语义保真度,明显优于传统的以失真为导向的JSCC方法。

英文摘要:

Conventional communication systems, including both separation-based coding and AI-driven joint source-channel coding (JSCC), are largely guided by Shannon's rate-distortion theory. However, relying on generic distortion metrics fails to capture complex human visual perception, often resulting in blurred or unrealistic reconstructions. In this paper, we propose Joint Source-Channel-Generation Coding (JSCGC), a novel paradigm that shifts the focus from deterministic reconstruction to probabilistic generation. JSCGC leverages a generative model at the receiver as a generator rather than a conventional decoder to parameterize the data distribution, enabling direct maximization of mutual information under channel constraints while controlling stochastic sampling to produce outputs residing on the authentic data manifold with high fidelity. We further derive a theoretical lower bound on the maximum semantic inconsistency with given transmitted mutual information, elucidating the fundamental limits of communication in controlling the generative process. Extensive experiments on image transmission demonstrate that JSCGC substantially improves perceptual quality and semantic fidelity, significantly outperforming conventional distortion-oriented JSCC methods.

补充信息

↑