基于 token 级记忆不对称性的微调扩散语言模型成员推断
Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry
- National University of Singapore(新加坡国立大学)
- Oregon State University(俄勒冈州立大学)
- Shanghai Jiao Tong University(上海交通大学)
- College of AI, Tsinghua University(清华大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对微调扩散语言模型的隐私风险,提出基于 token 级记忆不对称的 Q-Skew 指标用于成员推断,实验表明其优于基线且可助力 PII 提取,凸显 DLMs 系统性隐私评估的必要性。
AI中文摘要:
扩散语言模型(DLMs)作为替代自回归语言模型的建模范式,具备并行生成和双向上下文建模等优势,但其隐私风险尚未得到充分探索。本文通过对扩散训练动态的理论分析,发现了 token 级记忆不对称现象;基于该发现,提出了 Q-Skew 这一分位数加权偏度指标,用于对微调后的 DLMs 进行成员推断。在多个微调数据集和模型上开展的实验表明,所提方法优于现有基线,且 Q-Skew 还可助力 PII 提取等其他隐私侵犯行为。研究揭示了此前未被充分探索的隐私攻击面,强调需对 DLMs 开展系统性隐私评估。
英文摘要:
Diffusion language models (DLMs) have recently emerged as an alternative modeling paradigm to autoregressive LMs, offering advantages such as parallel generation and bidirectional context modeling. Despite growing interest in their generative capabilities, the privacy risks of DLMs remain underexplored. We identify a phenomenon termed token-level memorization asymmetry through theoretical analysis of diffusion training dynamics. Building on this finding, we propose Q-Skew, a quantile-weighted skewness-based indicator for membership inference on finetuned DLMs. Experiments across multiple fine-tuning datasets and models show that our method outperforms existing baselines. Moreover, we show that Q-Skew can also facilitate other privacy violations, such as PII extraction. Our findings reveal a previously underexplored privacy attack surface and highlight the need for systematic privacy evaluation of DLMs.