RAC:面向通信高效的分割式大语言模型推理的参考感知激活压缩
RAC: Reference-Aware Activation Compression for Communication-Efficient Split LLM Inference
浏览论文内容
中文总结 AI 辅助
该研究针对分割式LLM推理的通信瓶颈,提出参考感知激活压缩方法RAC,通过参考感知编解码器等技术降低通信开销,在多模型多链路对实验中验证了其时间性能提升与任务分数变化情况。
中文摘要 AI 辅助
大语言模型(LLM)智能体会反复处理长且隐私敏感的上下文,仅云部署会将用户数据暴露于可信端点之外,而全本地部署通常需要昂贵的硬件。分割推理提供了一种折中方案:将模型头部、尾部及工具在本地执行,中间层在云端执行,但该方案的本地-云端-本地路径会在每次调用时传输边界隐藏状态,形成关键的通信瓶颈。我们提出了RAC,一种参考感知编解码器,它为预装上行业务检索精确token的历史跨度,为同轮次预装下行业务复用重构的上行状态,并用轻量因果预测器生成特定边界的解码参考。RAC采用分组仿射对齐和带可选预装异常值的校准残差量化,同时发送方线路格式重构同步后续参考,离线校准兼顾质量和打包表示成本。在三个模型和9组评估的模型-链路对上,Raw到RAC的平均首token时间(TTFT)和单输出token时间(TPOT)比率分别为1.24-2.72倍和1.01-2.79倍,12项非困惑度任务的分数变化范围为-0.40到+2.50分。
英文摘要
Large language model (LLM) agents repeatedly process long, privacy-sensitive contexts, while cloud-only deployment exposes user data beyond the trusted endpoint and fully local deployment often requires costly hardware. Split inference offers a middle ground by executing the model head, tail, and tools locally and the middle layers in the cloud, but its local-cloud-local path transfers boundary hidden states at every invocation and creates a critical communication bottleneck. We present \system, a reference-aware codec that retrieves exact-token historical spans for prefill uplinks, reuses the reconstructed uplink state for same-round prefill downlinks, and generates boundary-specific decode references with lightweight causal predictors. RAC applies grouped affine alignment and calibrated residual quantization with optional prefill outliers, while sender-side wire-format reconstruction synchronizes subsequent references and offline calibration accounts for quality and packed representation costs. Across three models and nine evaluated model-link pairs, Raw-to-RAC mean time to first token (TTFT) and time per output token (TPOT) ratios are 1.24-2.72$\times$ and 1.01-2.79$\times$, while the 12 non-perplexity task-score changes range from $-0.40$ to $+2.50$ points.