AI 中文总结
AceSpec 是一种非对称边云协同框架,通过概率状态缓存、非对称通信协议与拉格朗日优化资源分配策略,实现 LLM 推理的通信高效性,最高获 3.52 倍吞吐量加速且带宽免疫性优异。
AI 中文摘要
将大语言模型(LLMs)部署在边缘设备通常依赖模型压缩或拆分推理,但压缩会降低推理能力,而拆分推理则面临严重的广域网(WAN)通信瓶颈。边云投机解码作为一种有前景的替代方案,利用边缘小模型生成待云验证的 token,然而在不稳定的 WAN 中,不可避免的预测拒绝会引发灾难性的流水线停顿和全网回滚,抵消协同增益。为解决该问题,我们提出 AceSpec,一种非对称边云协同框架。AceSpec 利用未饱和的边缘计算主动构建概率状态缓存,将全网流水线刷新转化为 O(1) 级别的本地内存查找;为节省带宽,采用非对称通信协议,上行传输最小的主链索引,下行传输紧凑的稀疏分布;此外,我们引入感知网络、拉格朗日优化的资源分配策略,动态最大化本地缓存命中率。评估表明,AceSpec 实现了最高 3.52 倍的吞吐量加速,且具有出色的带宽免疫性,即使在 50 Kbps 的严重受限 WAN 条件下也能维持接近峰值的推理性能。
英文摘要
Deploying Large Language Models (LLMs) on edge devices typically relies on model compression or split inference. However, compression degrades reasoning capabilities, while split inference suffers from severe Wide Area Network (WAN) communication bottlenecks. Edge-cloud speculative decoding emerges as a promising alternative, leveraging an edge small model to draft tokens for cloud verification. Yet, over volatile WANs, inevitable prediction rejections trigger catastrophic pipeline stalls and network-wide rollbacks, neutralizing collaborative gains. To overcome this, we propose AceSpec, an asymmetric edge-cloud collaborative framework. AceSpec utilizes un-saturated edge compute to proactively construct a probabilistic state cache, effectively transforming network-wide pipeline flushes into $\mathcal{O}(1)$ local memory lookups. To preserve bandwidth, it employs an asymmetric communication protocol that transmits minimal main-chain indices uplink and compact sparse distributions downlink. Furthermore, we introduce a network-aware, Lagrangian-optimized resource allocation strategy that dynamically maximizes the local cache hit rate. Evaluations demonstrate that AceSpec achieves up to a 3.52$\times$ throughput speedup and exhibits exceptional bandwidth immunity, sustaining near-peak inference performance even under severely constrained 50 Kbps WAN conditions.