面向任务的混合精度模型通信
Task-Oriented Communication with Hybrid-Precision Models
浏览论文内容
中文总结 AI 辅助
针对边缘推理中通信、计算和准确性难以平衡的问题,提出混合精度任务导向通信框架,通过二值化前端与全精度后端协同,结合特定方法和策略,在ImageNet数据集实验中表现优越,实现三者间最佳权衡。
中文摘要 AI 辅助
边缘推理通过在网络边缘部署模型来规避云路由延迟,已成为人工智能服务扩散的一种有前景的解决方案。现有边缘推理方法主要集中在协作推理或轻量级模型设计上,难以在传输效率、设备处理成本和推理准确性之间取得平衡。本文提出了一种面向任务的混合精度通信框架,在边缘设备上部署二值化前端通过正交频分复用信号提取和传输二值特征,边缘服务器上的全精度后端进行最终推理。为确保模型一致性,引入了针对分割推理的设备上二值化方法并开发了基于子载波特征校准的集成信道感知传输方案。此外,还开发了基于知识蒸馏的训练策略来优化端到端系统。在大规模ImageNet数据集上的大量实验证明了该混合系统的优越性,实现了通信效率、设备计算成本和推理准确性之间的最佳权衡,优于现有边缘推理解决方案。
英文摘要
Edge inference has emerged as a promising solution for the proliferation of artificial intelligence (AI) services by deploying models at the network edge to circumvent cloud-routing latency. Existing edge inference approaches mainly focused on either cooperative inference to reduce latency or lightweight model design to fit resource-constrained devices. These solutions often address the communication and computation challenges separately, and thus struggle to achieve a balanced trade-off among transmission efficiency, on-device processing cost, and inference accuracy. To bridge this gap, this paper proposes a hybrid-precision task-oriented communication framework for edge inference to holistically balance communication, on-device computation, and utility. In this framework, a binarized front-end is deployed on the edge device to extract and transmit binary features via orthogonal frequency-division multiplexing (OFDM) signals, while a full-precision back-end on the edge server performs the final inference. To ensure model consistency, we introduce an on-device binarization method tailored for split inference and develop an integrated channel-aware transmission scheme featuring subcarrier-based feature calibration. Furthermore, a knowledge distillation (KD)-based training strategy, supported by specialized gradient estimators, is developed to optimize the end-to-end system and inherit semantic knowledge from a full-precision teacher model. Extensive experiments on the large-scale ImageNet dataset demonstrate the superiority of the proposed hybrid system. Our analysis confirms that this design achieves an optimal trade-off among communication efficiency, on-device computational cost, and inference accuracy, outperforming existing edge inference solutions.