云边协同解码中的隐私泄露审计与缓解
Auditing and Mitigating Privacy Leakage in Cloud-Edge Collaborative Decoding
浏览论文内容
中文总结 AI 辅助
本文针对云边协同解码范式的隐私风险提出防御机制CoVeil,可降低数据泄露最多87.2%且准确率损失极小,改善了隐私-效用权衡。
中文摘要 AI 辅助
个性化辅助和专有文档分析等应用需要大型语言模型(LLM)基于私有数据生成输出。然而,强大的LLM通常无法部署在承载私有数据的资源受限设备上,将私有数据上传至云端托管的LLM会暴露敏感信息。近期研究提出了云边协同解码范式来解决这一矛盾:私有数据保留在边缘端,小型语言模型(SLM)生成下一个token的分布,该分布与仅基于公共数据运行的云端LLM的预测相融合。本文使用构建的问答数据集,通过一种新颖的评估框架系统分析了该范式的隐私风险,结果显示此类协同会暴露大量私有上下文信息。为解决此类隐私泄露问题,我们提出了防御机制CoVeil,该机制会在解码阶段动态优化传输信号,以抑制泄露同时保持协同质量。大量评估表明,与现有基线相比,CoVeil可将数据泄露最多降低87.2%,且准确率损失极小,持续改善了隐私-效用权衡。
英文摘要
Applications such as personalized assistance and proprietary document analysis require large language models (LLMs) to generate outputs from private data. Yet powerful LLMs typically cannot be deployed on the resource-constrained devices where private data resides, and uploading private data to cloud-hosted LLMs exposes sensitive information. Recent work addresses this tension with a cloud-edge collaborative decoding paradigm, where private data are kept on the edge with a small language model (SLM) producing next-token distributions, which are fused with predictions from a cloud LLM operating solely on public data. In this paper, we systematically analyze the privacy risks of such a paradigm with a novel evaluation framework using constructed QA datasets, which show that such collaboration can expose substantial private-context information. To address such privacy leakage, we propose CoVeil, a defense mechanism which dynamically optimizes transmitted signals to suppress leakage during decoding time while preserving the collaborative quality. Extensive evaluations demonstrate that CoVeil consistently improves the privacy-utility trade-off over existing baselines by reducing data leakage by up to 87.2%, with minimal accuracy loss.
发表机构
- The Hong Kong Polytechnic University(香港理工大学)
- Institute for AI Industry Research, Tsinghua University(清华大学人工智能产业研究院)
- School of Software, Tsinghua University(清华大学软件学院)
机构由 AI 辅助整理,请以论文原文为准。