AI 中文总结
针对AI-RAN中大型AI模型的高效推理挑战,提出剪枝感知多集群协同推理框架,通过联合优化实现更优的推理精度与资源效率。
AI 中文摘要
大型人工智能模型(LAIMs)规模不断扩大,计算需求日益增长,这给资源受限的分布式环境中的高效推理带来了重大挑战。本文提出一种多集群LAIM协同推理框架,配备多个图形处理单元(GPUs)的边缘服务器协调多个用户集群协同执行推理任务。每个集群内的设备从不同视角采集数据,采用轻量级设备端LAIM提取局部特征,这些特征随后被传输至边缘服务器,经聚合与融合生成更准确的推理结果。为揭示模型剪枝与协同推理性能间的基本权衡,我们构建了一个理论框架,利用率失真理论和偏信息分解刻画剪枝率与设备贡献的影响。基于此分析,我们建立联合优化问题,确定模型剪枝率、任务调度策略、带宽分配及传输功率,目标是在满足延迟、能耗和服务器容量约束的同时最小化推理失真。大量仿真结果表明,所提框架显著优于现有基准方案,在多集群边缘智能网络中实现了更优的推理精度与资源效率。
英文摘要
The increasing scale and computational demands of large artificial intelligence models (LAIMs) present significant challenges for efficient inference in resource-constrained distributed environments. In this paper, we propose a multi-cluster LAIM co-inference framework, where an edge server equipped with multiple graphics processing units (GPUs) coordinates multiple user clusters to execute inference tasks collaboratively. Within each cluster, devices capture data from diverse perspectives and employ lightweight on-device LAIMs to extract local features. These features are then transmitted to the edge server, where they are aggregated and fused to generate a more accurate inference outcome. To reveal the fundamental trade-off between model pruning and collaborative inference performance, we develop a theoretical framework that characterizes the impact of pruning ratios and device contributions using rate-distortion theory and partial information decomposition. Based on this analysis, we formulate a joint optimization problem that determines the model pruning ratio, the task scheduling strategy, the bandwidth allocation, and the transmission power, with the goal of minimizing the inference distortion while satisfying the constraints of latency, energy consumption, and server capacity. Extensive simulation results demonstrate that the proposed framework significantly outperforms existing benchmark schemes, achieving superior inference accuracy and resource efficiency in multi-cluster edge intelligence networks.