arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29618cs.ITmath.IT

多接入投机推理:上行链路还是下行链路?

Multi-Access Speculative Inference: Uplink or Downlink?

Chang Cai, Kaibin Huang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多接入投机推理,提出自适应选择上行/下行链路的通信模式,结合草拟长度控制与功率分配优化,缓解通信瓶颈,在Qwen2.5与DeepSeek-R1模型对上显著提升令牌吞吐量。

中文摘要 AI 辅助

多接入投机推理(Multi-SPIN)将投机推理(SPIN)扩展至多设备边缘网络,以加速协作式令牌生成。它允许设备端小型语言模型(SLM)为各自的生成任务自回归地草拟多个令牌,而边缘服务器端的大型语言模型(LLM)则并行验证这些令牌。主要的通信开销产生于服务器拒绝某一个草拟令牌时,此时采样修正令牌需要同时访问SLM输出的草拟分布和LLM输出的完整令牌词汇表上的目标分布。现有设计通常通过上传草拟分布在服务器端执行修正,但传输词汇表范围的分布会造成严重的上行链路(UL)瓶颈。另一种方案是在设备端执行修正,通过下载目标分布来利用下行链路(DL)的高传输速率。受此启发,本文将通信模式选择作为Multi-SPIN的新设计维度,具体而言,每个设备可自适应地在UL和DL模式间切换,以平衡UL瓶颈与共享DL资源约束,从而减轻整体通信负担。本文构建了一个令牌总吞吐量最大化问题,该问题同时考虑模式选择、草拟长度控制和功率分配;对于模式选择,本文揭示了一种简单的最优结构,可对UL设备数量进行高效搜索,并相应优化对应的发射功率;对于草拟长度控制,本文开发了一种贪心搜索算法,可使设备特定的草拟长度适配异构计算与通信能力。在Qwen2.5和DeepSeek-R1模型对上的实验结果表明,所提出的框架显著提升了令牌吞吐量。

英文摘要

Multi-access speculative inference (Multi-SPIN) extends SPIN to multi-device edge networks to accelerate cooperative token generation. It allows on-device small language models (SLMs) to autoregressively draft multiple tokens for individual generation tasks, while an edge-server large language model (LLM) verifies them in parallel. The major communication overhead arises when a drafted token is rejected by the server, in which case sampling the correction token requires access to both the SLM-output draft distribution and the LLM-output target distribution over the full token vocabulary. Existing designs typically perform correction at the server by uploading the draft distribution, but transmitting a vocabulary-wide distribution creates a critical uplink (UL) bottleneck. Alternatively, the correction can be performed at the device by downloading the target distribution, leveraging the high transmission rates available on the downlink (DL). Motivated by this insight, we introduce communication-mode selection as a new design dimension for Multi-SPIN. Specifically, each device can adaptively switch between the UL and DL modes to balance the UL bottleneck against the shared DL resource constraint, thereby relieving the overall communication burden. We formulate a sum-token-goodput maximization problem that jointly accounts for mode selection, draft-length control, and power allocation. For mode selection, we reveal a simple optimal structure that enables efficient search over the number of UL devices, with the corresponding transmit powers optimized accordingly. For draft-length control, we develop a greedy-search algorithm that adapts device-specific draft lengths to heterogeneous computation and communication capabilities. Experimental results on Qwen2.5 and DeepSeek-R1 model pairs demonstrate that the proposed framework significantly improves token goodput.

发表机构

  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑