发表机构
Nanyang Technological University; National University of Singapore(南洋理工大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对传统GPU上可验证的差分隐私训练问题,提出利用CPU可信执行环境与不可信GPU协同的框架,通过概率检查验证梯度DP执行,以低开销实现高概率检测偏离并保持效用。
AI 中文摘要
机器学习的广泛应用带来了日益增长的政策和监管需求,要求保护敏感训练数据,差分隐私(DP)已成为一种关键机制。然而,一个研究较少的问题是,如何认证训练过程中DP的忠实执行:外部验证者应能在不访问私有训练数据的情况下,检查发布的模型是否经过适当的DP保护训练。现有的密码学方法,如零知识证明,提供了强有力的保证,但往往带来高昂的开销,在某些情况下甚至高出数个数量级。可信执行环境(TEE)提供了一种更高效的替代方案,但训练和微调大型语言模型所需的多GPU TEE支持仅限于最新平台,在传统GPU上缺失或效率低下。为解决这一问题,我们提出了一个实用的框架,利用CPU侧TEE配合不可信的GPU实现可验证的DP训练。我们的设计解决了一个基本的效率-安全矛盾:完全在CPU TEE内训练过于缓慢,而无限制的GPU卸载可能允许恶意偏离DP。因此,我们将昂贵的梯度计算卸载到GPU,同时使用CPU TEE通过概率检查高效验证梯度上DP的正确执行。我们的框架能以高概率检测频繁的完全偏离DP行为;对于本工作中评估的以效用为导向的伪造梯度攻击,稀疏偏离仅提供有限的效用收益,且未显示可测量的额外成员泄露。实验进一步表明,我们的方法几乎实现了“免费午餐”:与基于标准GPU的DP训练相比,仅带来适度的开销,同时有效约束了声称的DP执行中的恶意偏离。
英文摘要
Wide adoption of machine learning has created growing policy and regulatory demand for protecting sensitive training data, with differential privacy (DP) emerging as a key mechanism. Yet a less-studied problem is how to certify the faithful execution of DP during training: an external verifier should be able to check that a released model was trained with proper DP protection, without accessing the private training data. Existing cryptographic approaches, such as zero-knowledge proofs, provide strong guarantees but often incur prohibitive overhead, in some cases by orders of magnitude. Trusted Execution Environments (TEEs) offer a more efficient alternative, but the multi-GPU TEE support needed for training and fine-tuning large language models remains limited to recent platforms and is absent or inefficient on legacy GPUs. To address this, we propose a practical framework for verifiable DP training using CPU-side TEEs together with untrusted GPUs. Our design addresses a fundamental efficiency-security tension: training entirely inside a CPU TEE is too slow, while unrestricted GPU offloading can allow malicious deviations from DP. We therefore offload expensive gradient computation to GPUs, while using the CPU TEE to efficiently verify the correct enforcement of DP on gradients through probabilistic checking. Our framework detects frequent full deviations from DP with high probability; for the utility-oriented forged-gradient attacks evaluated in this work, sparse deviations provide limited utility benefit and show no measurable additional membership leakage. Experiments further show that our approach nearly achieves a ``free lunch'': it incurs only modest overhead compared with standard GPU-based DP training, while effectively constraining malicious deviations from the claimed DP execution.