arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向泛化的语义与目标导向通信的基础模型

Foundation Models for Generalizable Semantic and Goal-Oriented Communication

Boliang Liu, Wint Yi Poe, Riccardo Trivisonno, Giuseppe Caire

arXiv 2609.07853首次发表:更新:

发表机构

Technical University of Berlin; Huawei Technologies Düsseldorf GmbH, Munich Research Center(柏林工业大学; 华为技术杜塞尔多夫有限公司慕尼黑研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出FMSGOC框架,利用视觉-语言基础模型选择稀疏语义锚点,结合扩散模型重建,实现低比特率下高泛化与保真度的语义通信。

AI 中文摘要

语义与目标导向通信在6G中日益受到研究关注,但在严格速率预算下,对未见数据的泛化能力仍是关键弱点。许多现有系统过度拟合训练数据,在极低比特率下性能急剧下降,因为它们试图压缩整个信号。我们提出了基础模型引导的语义与目标导向通信(FMSGOC),该框架利用广泛的视觉-语言基础模型先验来缓解过拟合。它通过将比特集中在稀疏的、目标对齐的锚点上,并依赖生成式基础模型先验来重建掩蔽区域,从而进一步提高速率效率。通过将发送内容与重建方式解耦,视觉-语言基础模型选择并传输一组稀疏的语义锚点,而预训练的扩散模型(针对掩蔽补全进行微调)在接收端重建图像。在我们的实验中,FMSGOC达到0.039比特每像素(BPP),在CIFAR-10上保持高语义保真度(余弦相似度0.87-0.90),在未见输入上保持鲁棒性(ImageNet上0.83-0.86),并显示出良好的感知相似性(CIFAR-10/ImageNet分别为0.1278/0.1558),在较低比特率下优于强端到端基线。

英文摘要

Semantic and goal-oriented communication is increasingly studied for 6G, but generalization beyond seen data remains a key weakness under tight rate budgets. Many existing systems overfit their training data and degrade sharply at very low bit rates because they attempt to compress the entire signal. We introduce Foundation Model-Guided Semantic and Goal-Oriented Communication (FMSGOC), a framework that uses broad visual-linguistic Foundation Model priors to mitigate overfitting. It further improves rate efficiency by concentrating bits on sparse, goal-aligned anchors and relying on generative foundation-model priors to reconstruct the masked regions. By decoupling what to send from how to reconstruct, a vision-language foundation model selects and transmits a sparse set of semantic anchors, while a pretrained diffusion model, fine-tuned for masked completion, reconstructs the image at the receiver. In our experiments, FMSGOC reaches 0.039 bits per pixel (BPP), maintains high semantic fidelity (cosine similarity 0.87-0.90 on CIFAR-10), remains robust on previously unseen inputs (0.83-0.86 on ImageNet), and shows good perceptual similarity (0.1278/0.1558, CIFAR-10/ImageNet), outperforming strong end-to-end baselines at lower bit rates.

Comments6 pages, IEEE ICC 2026

Journal ref2026 IEEE International Conference on Communications (ICC), pp. 1-6, 2026

DOI:10.1109/ICC59461.2026.11588119

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑