arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型的高效且隐私感知的边缘云协作推理

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models

Cheng Li, Jiexiong Liu, Yixuan Chen, Yi Li

arXiv 2607.13093首次发表:更新:

发表机构

KunlunMeta(昆仑元)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究大语言模型推理难题,提出基于端点认证键值缓存的边缘云协作推理框架,本地端点与云端分工协作,经优化支持多种设备,评估显示该框架能降延迟和载荷,性能与全云端推理相当。

AI 中文摘要

设备端的大语言模型推理面临响应延迟、硬件资源有限和用户隐私这三个难题。全云端推理虽有强大计算能力,但会暴露用户提示和对话数据,而独立的设备端推理对大多数消费和嵌入式边缘设备不可行。本文提出了一个以隐私为中心的边缘云协作大语言模型推理框架,该框架基于端点认证的键值缓存构建。本地端点处理输入预处理、嵌入计算、自适应特征优化、键值缓存认证、推测性解码和低维模型头部计算,而云端进行认证解码器推理、键值缓存管理、令牌验证和高维词汇投影。端点融合部分输出,应用语言自适应掩码并采样目标令牌。所有传输的数据和截断的对数概率都进行量化并通过AES - GCM加密以保护隐私,核心轻量级模块、草稿参数和缓存访问策略都保存在本地以避免泄露。该框架通过优化的流、批处理和量化的ONNX部署支持包括仅CPU、配备GPU和嵌入式设备在内的异构设备。评估表明,该框架与基线分割推理相比,每个令牌的延迟最多降低46.1%,下行链路有效载荷最多降低67.4%,同时保持与全云端推理相当的性能。

英文摘要

On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers strong computing power but exposes user prompts and dialogue data, while standalone on-device inference is unfeasible for most consumer and embedded edge devices. This paper presents a privacy-centric edge-cloud collaborative LLM inference framework built on endpoint-authenticated KV cache. Local endpoints handle input preprocessing, embedding computation, adaptive feature optimization, KV cache authentication, speculative decoding and low-dimensional model head calculation, while the cloud conducts authenticated decoder inference, KV cache management, token verification and high-dimensional vocabulary projection. Endpoints fuse partial outputs, apply language-adaptive masking and sample target tokens. All transmitted data and truncated logits are quantized and AES-GCM encrypted for privacy, with core lightweight modules, draft parameters and cache access policies kept local to avoid leakage. The framework supports heterogeneous devices including CPU-only, GPU-equipped and embedded devices via optimized streaming, batching and quantized ONNX deployment. Evaluations demonstrate that the framework reduces per-token latency by up to 46.1\% and downlink payloads by up to 67.4\% over baseline split inference, retaining comparable performance to full cloud inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑