发表机构
Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出X-CoSD及增强版X-CoSD-E,通过混合重采样和服务器重采样-设备验证机制,实现跨词表场景下无损且通信高效的协作式投机解码,显著提升生成速度并保持生成质量。
AI 中文摘要
本文研究了协作式投机解码(CoSD),这是一种分布式大语言模型(LLM)推理框架,其中设备端的小语言模型(SLM)负责生成候选词元,而服务器端的LLM负责验证这些词元。现有的CoSD方法假设SLM和LLM共享词表,并且由于残差重采样需要在用户设备与边缘服务器之间交换词元分布,因此会产生大量的通信负载。为了解决这些限制,我们提出了跨词表CoSD(X-CoSD),这是一种适用于异构SLM-LLM词表的无损且通信高效的CoSD框架。X-CoSD基于混合重采样(HR),该技术将残差重采样在设备上的公共词表区域和服务器上的LLM专属区域之间进行拆分,从而仅需对公共词表区域进行分布传输。我们进一步提出了X-CoSD-E,这是一种基于服务器重采样与设备验证(SR-DV)的增强变体,其中服务器仅发送从服务器LLM中采样的替换候选词元及其对应概率,供设备进行本地验证。我们证明了X-CoSD和X-CoSD-E都能保持服务器LLM的分布,并且实验表明,它们在保持与服务器LLM相当的生成质量的同时,显著提高了词元生成速度。
英文摘要
This paper investigates collaborative speculative decoding (CoSD), a distributed large language model (LLM) inference framework in which an on-device small language model (SLM) drafts candidate tokens and a server LLM verifies them. Existing CoSD methods assume a shared vocabulary between the SLM and the LLM and incur substantial communication load because residual resampling requires token distribution exchange between the user device and the edge server. To address these limitations, we propose cross-vocabulary CoSD (X-CoSD), a lossless and communication-efficient CoSD framework for heterogeneous SLM-LLM vocabularies. X-CoSD is built on hybrid resampling (HR), which splits residual resampling across the common-vocabulary region on the device and the LLM-only region on the server, so that distribution transmission is required only for the common-vocabulary region. We further propose X-CoSD-E, an enhanced variant based on server resampling with device verification (SR-DV), in which the server sends only replacement candidates sampled from the server LLM and their corresponding probabilities for local verification at the device. We prove that both X-CoSD and X-CoSD-E preserve the server LLM distribution, and experiments show that they significantly improve token generation speed while maintaining generation quality comparable to that of the server LLM.