发表机构
State Key Laboratory of Advanced Rail Autonomous Operation; School of Electronic and Information Engineering, Beijing Jiaotong University(轨道交通运行控制系统国家实验室; 北京交通大学电子信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HintKD通过将教师特征压缩为离散提示索引,将云边推理通信负载降至O(1),在保持精度的同时实现带宽受限下的高效协同推理。
AI 中文摘要
边缘云与用户设备(UE)之间的协同推理是在6G网络中部署大型深度神经网络(DNNs)的一种有前景的范式。然而,现有的云边推理和蒸馏方案通常需要实时传输高维中间特征或软概率向量,这给带宽受限的无线链路带来了巨大的通信负担。为应对这一挑战,我们提出了HintKD,一种用于云边推理的带宽受限蒸馏框架。其核心思想是将教师的指导压缩为紧凑的离散提示,而非直接传输原始特征。具体而言,首先使用最大条件互信息(MCMI)目标训练云侧教师,以保留信息丰富的类内变异。然后通过可微向量量化将其潜在表示映射到紧凑的码本。在部署期间,云仅传输一个提示索引,用户设备使用基于FiLM的轻量级适配器根据接收到的码字调制学生特征。在DeepSense 6G波束预测任务上的实验表明,HintKD在保持有竞争力的精度的同时,将推理侧通信负载从特征传输的O(D_t)或逻辑传输的O(C)降低到离散提示传输的O(1)。此外,在MNIST数据集上的实验进一步验证了所提蒸馏方法在不同数据集和任务上的通用性。这些结果表明,HintKD通过用紧凑的离散提示索引替代高维特征或逻辑传输,实现了良好的精度-带宽权衡。
英文摘要
Collaborative inference between the edge cloud and user equipment (UE) is a promising paradigm for deploying large deep neural networks (DNNs) in 6G networks. However, existing cloud-edge inference and distillation schemes often require the real-time transmission of high-dimensional intermediate features or soft probability vectors, which imposes a substantial communication burden on bandwidth-limited wireless links. To address this challenge, we propose HintKD, a bandwidth-constrained distillation framework for cloud-edge inference. The core idea is to compress the teacher's guidance into compact discrete hints rather than transmit raw features directly. Specifically, a cloud-side teacher is first trained with a maximum conditional mutual information (MCMI) objective to preserve informative intra-class variations. Its latent representation is then mapped to a compact codebook through differentiable vector quantization. During deployment, the cloud transmits only a hint index, and the UE uses a lightweight FiLM-based adapter to modulate student features according to the received codeword. Experiments on the DeepSense 6G beam prediction task show that HintKD preserves competitive accuracy while reducing the inference-side communication payload from O(D_t) for feature transmission or O(C) for logit transmission to O(1) for discrete hint transmission. In addition, experiments on the MNIST dataset further verify the generality of the proposed distillation method across different datasets and tasks. These results show that HintKD achieves a favorable accuracy-bandwidth trade-off by replacing high-dimensional feature or logit transmission with a compact discrete hint index.