发表机构
The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出文本驱动图像语义通信框架TISC,通过树结构属性语义提取(TSASE)和初始噪声优化(INO)机制解决现有方法的语义损失与重建忠实度不足问题,经多数据集实验验证了其有效性。
AI 中文摘要
生成式图像语义通信将图像转换为文本描述,再在接收端通过基于扩散的生成模型进行文本到图像的重建,该范式因带宽成本极低而受到广泛关注。然而,现有方法在发送端的图像到文本(I2T)语义提取和接收端的文本到图像(T2I)语义重建两个关键瓶颈上仍存在问题:(i)I2T中的语义损失与失真,整体图像描述可能忽略细粒度对象属性和空间位置信息,导致生成的文本偏离原始图像语义;(ii)T2I中的语义忠实度不足,即使使用语义忠实的文本描述,不同的初始噪声设置也可能使基于扩散的重建生成与原始图像语义一致性不同的图像。这些问题共同限制了图像重建的语义忠实度。为解决这些问题,本文提出TISC,一种专为忠实重建设计的文本驱动图像语义通信框架。TISC包含两个关键设计:(1)树结构属性语义提取(TSASE),将语义提取分解为全局场景、背景和对象级属性描述,覆盖每个检测对象的空间位置、形状/姿态、颜色、材质及其他物理属性;(2)初始噪声优化(INO)机制,发送端根据综合相似度分数选择初始噪声种子,该分数同时考虑视觉和语义一致性。在多个数据集上的实验表明,TSASE提升了对象位置恢复和语义描述的忠实度,而INO参数研究验证了所采用的噪声选择配置的有效性。
英文摘要
Generative image semantic communication converts an image into a text description and then performs text-to-image reconstruction at the receiver via diffusion-based generative models. This paradigm has attracted broad attention due to its extremely low bandwidth cost. However, existing methods still face two critical bottlenecks across image-to-text (I2T) semantic extraction at the transmitter and text-to-image (T2I) semantic reconstruction at the receiver: (i) semantic loss and distortion in I2T, where holistic image descriptions may omit fine-grained object attributes and spatial-position information, causing the generated text to deviate from the original image semantics; and (ii) insufficient semantic faithfulness in T2I, where even with the same semantically faithful text description, different initial noise settings may lead diffusion-based reconstruction to produce images with different levels of semantic consistency with the original image. These issues jointly limit the semantic faithfulness of image reconstruction. To address them, we propose TISC, a text-driven image semantic communication framework tailored for faithful reconstruction. TISC incorporates two key designs: (1) Tree-Structured Attribute Semantic Extraction (TSASE), which decomposes semantic extraction into global scene, background, and object-level attribute descriptions, covering spatial position, shape/pose, color, material, and other physical attributes for each detected object; and (2) an Initial Noise Optimization (INO) mechanism, which selects an initial noise seed at the transmitter according to a comprehensive similarity score that jointly considers visual and semantic consistency. Experiments on multiple datasets show that TSASE improves object-position recovery and semantic description faithfulness, while the INO parameter study supports the adopted configuration for noise selection.