arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13405eess.SP

Krum启发的中心教师选择与残差通道瓶颈用于高效DeepSC

Krum-Inspired Central Teacher Selection and Residual Channel Bottlenecks for Efficient DeepSC

Rami Eid, Mostafa Jammoul, Omar Kaaki, Maria Slim, Mariette Awad, Hadi Sarieddeen

首次发表
浏览论文内容

中文总结 AI 辅助

针对边缘部署的语义通信模型压缩问题,提出残差通道瓶颈与Krum启发中心教师选择方法,在EuroParl基准上以更少参数和延迟恢复教师大部分性能。

中文摘要 AI 辅助

在边缘设备上部署基于Transformer的语义通信模型,需要在信道变化条件下进行压缩并保持语义保真度。我们研究了一个由多教师知识蒸馏训练的压缩深度语义通信(DeepSC)学生模型,并融合了两个思想:(i)残差通道瓶颈,将传输表示拆分为基础流和残差流,并采用不等功率分配;(ii)一种Krum启发的、类medoid的中心性准则,从五模型集成中选择单个中心教师,应用于logit层和中间特征层,并与特征主导的蒸馏损失相结合。在EuroParl基准上,一个两层学生模型在加性高斯白噪声下恢复了四层教师约98%的双语评估替代(BLEU)-1分数,恢复了其93%的BLEU-4和96%的句子-BERT(SBERT)分数;在瑞利衰落信道下,分别恢复了教师BLEU-1、BLEU-4和SBERT的约89%、77%和87%,同时将非嵌入参数减少了1.33倍,单教师推理延迟降低了1.79倍(相对于此处用作上界参考而非部署基线的五教师集成,延迟降低了9.0倍)。在三个随机种子上的受控消融实验表明,残差瓶颈带来了约+8.4%的BLEU-1(+9.6%的SBERT)提升,而基于中心性的选择相对于均值聚合带来了+3.0%(+2.9%的SBERT)的提升,两者贡献互补。解码器模式消融显示,仅基础流即可恢复完整BLEU-1的三分之二左右,而仅残差流则完全崩溃,这支持了残差作为基础流之上细化模块的作用。

英文摘要

Deploying transformer-based semantic communication models on edge devices requires compression that preserves semantic fidelity under channel variability. We study a compressed deep semantic communication (DeepSC) student trained by multi-teacher knowledge distillation and combine two ideas: (i) a residual channel bottleneck that splits the transmitted representation into a base stream and a residual stream with unequal power allocation, and (ii) a Krum-inspired, medoid-style centrality criterion that selects a single central teacher from a five-model ensemble, applied at the logit and intermediate-feature levels and combined with a feature-dominant distillation loss. On the EuroParl benchmark, a two-layer student recovers about 98% of the four-layer teacher's bilingual evaluation understudy (BLEU)-1 under additive white Gaussian noise, with 93% of its BLEU-4 and 96% of its sentence-BERT (SBERT) score, and about 89%, 77%, and 87% of the teacher's BLEU-1, BLEU-4, and SBERT, respectively, under Rayleigh fading, while reducing non-embedding parameters by 1.33x and single-teacher inference latency by 1.79x (9.0x relative to a five-teacher ensemble used here as an upper-bound reference, not a deployment baseline). Controlled ablations over three seeds indicate complementary contributions of about +8.4% BLEU-1 (+9.6% SBERT) from the residual bottleneck and +3.0% (+2.9% SBERT) from centrality-based selection over mean aggregation. A decoder-mode ablation shows that the base stream alone recovers about two-thirds of the full BLEU-1 while the residual stream alone collapses, supporting the role of the residual as a refinement on top of the base.

发表机构

  • American University of Beirut(贝鲁特美国大学)

机构由 AI 辅助整理,请以论文原文为准。

↑