arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40335cs.LG

在DP-SGD下的私有设置中,权重共享对仅解码器LLM仍然有益吗?

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Razan El Mais, Ali Chehab, Ibrahim Issa, Razane Tajeddine

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨DP-SGD下权重共享对仅解码器LLM的影响,发现未共享嵌入在效用和内存效率上均优于权重共享,为隐私保护训练提供了更优的架构选择。

中文摘要 AI 辅助

差分隐私随机梯度下降(DP-SGD)是隐私保护大型语言模型(LLM)微调的主要方法。许多仅解码器LLM在输入和输出嵌入之间采用权重共享,这一设计最初是为了在非私有设置中提高参数效率和改善语言建模性能。然而,在差分隐私训练下权重共享的影响在很大程度上尚未被探索。在本工作中,我们使用GPT2和DistilGPT2作为代表性的仅解码器架构,研究权重共享在DP设置中的作用。有趣的是,我们发现未共享嵌入在DP-SGD下始终优于权重共享模型,在SST-2、QNLI和QQP上实现了高达4.74个百分点的准确率提升。除了改进的效用,未共享嵌入还使得DP-SGD能够使用内存高效的幽灵裁剪。相比之下,权重共享引入了共享参数交互,使标准幽灵范数计算复杂化,并在很大程度上抵消了其计算优势。因此,未共享模型在保持幽灵裁剪优势的同时,内存使用量降低了60%以上。我们的结果表明,未共享嵌入为仅解码器LLM的差分隐私训练提供了更有效和可扩展的设计,并强调了在隐私保护设置中重新审视标准LLM架构选择的必要性。

英文摘要

Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points in accuracy on SST-2, QNLI, and QQP. Beyond improved utility, untying embeddings enables the use of memory-efficient ghost clipping for DP-SGD. By contrast, weight tying introduces shared-parameter interactions that complicate standard ghost norm computation and largely negate its computational advantages. As a result, untied models achieve over 60% lower memory usage while preserving the benefits of ghost clipping. Our results indicate that untied embeddings provide a more effective and scalable design for differentially private training of decoder-only LLMs and highlight the need to revisit standard LLM architectural choices in the privacy-preserving setting.

发表机构

  • American University of Beirut(贝鲁特美国大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑