发表机构
IIT Bombay(印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SPADE是一种集成推测解码的分布式边缘云推理框架,通过边缘草稿模型与云端验证模型的协作,减少76%云端模型调用且不损失准确率,降低了LLM部署的成本与推理时间。
AI 中文摘要
大语言模型(LLMs)在自然语言理解与生成领域取得了显著成功,但其部署受限于高昂的计算需求。直接在边缘部署较小的LLMs可规避该问题,但准确率会下降;在云端部署较大的LLMs可保持性能,但代价是昂贵的单令牌计算成本。本文提出了一种分布式推理框架SPADE,它在边缘与云端间集成了推测解码(SD):部署在边缘的紧凑草稿模型可快速生成候选令牌,云端的大型验证模型并行验证这些令牌;仅当验证模型拒绝候选令牌时才触发修正,从而大幅减少云端查询次数。该即插即用设计将大部分计算转移至边缘,显著降低推理时间与云端成本,且无需任何再训练即可保持大型模型的准确率。在SpecBench与CNN/Dailymail数据集上的多项自然语言处理任务实验结果显示,与完整模型相比,SPADE可减少76%的云端模型调用,且准确率无损失。
英文摘要
Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is constrained by high computational demands. Deploying smaller LLMs directly on the edge can circumvent this, but with degraded accuracy. Deploying smaller cloud-based big LLMs preserves performance, but at the cost of expensive per-token computation. We present a distributed inference framework, \our{}, that integrates speculative decoding (SD) across edge and cloud. A compact draft model deployed on the edge generates candidate tokens rapidly, and a large verifier model on the cloud validates these tokens in parallel. Accepted tokens are retained, while only rejections trigger verifier correction, substantially reducing the number of cloud queries. Our plug-and-play design shifts the bulk of computation to the edge, significantly lowers inference time and cloud cost, and preserves the accuracy of the big model without any retraining requirement. Our approach demonstrates a practical path toward scalable, cost-efficient, and accurate deployment of LLMs in real-world environments. Experimental results across multiple Natural Language Processing tasks using SpecBench and CNN/Dailymail datasets demonstrate that \our{} reduces the cloud model calls by $76\%$ with zero loss in accuracy as compared to the full model.