FlowSem:语义通信中自适应无线图像传输的流匹配方法
FlowSem: Flow Matching for Adaptive Wireless Image Transmission in Semantic Communication
浏览论文内容
中文总结 AI 辅助
FlowSem是一种两阶段流匹配语义通信框架,通过自适应DeepJSCC与条件流匹配结合,在Cityscapes数据集的不同信道条件下,比扩散基线FID低60%且质量延迟权衡更优。
中文摘要 AI 辅助
在信道条件差、带宽约束严格的情况下,无线图像传输面临挑战,因为接收端需要同时保留像素级保真度和有意义的视觉结构。传统的分离式系统,如采用低密度奇偶校验编码的更好便携式图形格式(BPG+LDPC),可能会出现 cliff-effect(悬崖效应)行为。深度联合信源信道编码(DeepJSCC)虽然在信道条件恶化时性能会平缓下降,但在强压缩和严重信道失真下,其重建结果可能会丢失精细细节。为解决这一局限,本文提出了一种基于流匹配的两阶段语义通信框架,命名为FlowSem。第一阶段,采用信噪比(SNR)自适应的DeepJSCC模型将源图像映射为信道符号,并在接收端生成粗重建结果;第二阶段,采用条件流匹配模型,基于DeepJSCC的重建结果和信道SNR,从高斯噪声中生成最终图像。在加性高斯白噪声(AWGN)和瑞利衰落信道下,使用固定的信道符号预算,在Cityscapes数据集上对所提框架进行评估。基线方法包括速率匹配的BPG+LDPC系统、DeepJSCC、去噪扩散概率模型(DDPM)和去噪扩散隐式模型(DDIM)。结果表明,在不同信道条件下,FlowSem在像素级保真度上具有竞争力,且相比所考虑的生成式基线,其结构和感知重建质量有所提升;在低SNR下,FlowSem的Fréchet Inception Distance(FID)比扩散基线低达60%;此外,FlowSem仅需几步ODE积分即可达到高重建质量,相比标准DDPM和加速DDIM采样,具有良好的质量-延迟权衡。
英文摘要
Wireless image transmission becomes challenging under poor channel conditions and stringent bandwidth constraints, as the receiver needs to preserve both pixel-level fidelity and meaningful visual structure. Classical separation-based systems, such as better portable graphics with low-density parity-check coding (BPG+LDPC), may suffer from cliff-effect behavior. While deep joint source-channel coding (DeepJSCC) provides graceful degradation as channel conditions worsen, its reconstructions may lose fine details under strong compression and severe channel distortion. To address this limitation, this paper proposes a two-stage flow matching-based semantic communication framework, termed FlowSem. In the first stage, a signal-to-noise ratio (SNR)-adaptive DeepJSCC model maps the source image into channel symbols and produces a coarse reconstruction at the receiver. In the second stage, a conditional flow matching model generates the final image from Gaussian noise conditioned on the DeepJSCC reconstruction and channel SNR. The proposed framework is evaluated on the Cityscapes dataset under additive white Gaussian noise (AWGN) and Rayleigh fading channels using a fixed channel-symbol budget. The baselines include a rate-matched BPG+LDPC system, DeepJSCC, a denoising diffusion probabilistic model (DDPM), and a denoising diffusion implicit model (DDIM). Results show that FlowSem achieves competitive pixel-level fidelity and improved structural and perceptual reconstruction quality over the considered generative baselines across different channel conditions. FlowSem provides up to 60% lower Fréchet Inception Distance (FID) than the diffusion baselines at low SNRs. Moreover, FlowSem reaches high reconstruction quality using only a few ODE integration steps, providing a favorable quality-latency tradeoff compared with standard DDPM and accelerated DDIM sampling.