JUMP:微调扩散语言模型的单通道成员推理
JUMP: Efficient Membership Inference on Fine-Tuned Diffusion Language Models
浏览论文内容
中文总结 AI 辅助
研究微调离散扩散语言模型的成员推理攻击,提出JUMP单通道评分攻击,利用dLLMs任意顺序可解码性和并行可解码性,在六个MIMIR域上的微调LLaDA - 8B - Base上提升了平均ROC - AUC并改善低FPR检测,且减少查询次数。
中文摘要 AI 辅助
成员推理攻击(MIAs)用于测试候选示例是否出现在模型的训练数据中。我们研究了针对微调离散扩散语言模型(dLLMs)的MIAs,其中成员资格指包含在目标模型的微调集中。与自回归语言模型不同,dLLMs允许攻击者选择任意掩码集并并行获取所有掩码位置的令牌分布。先前的dLLM攻击SAMA通过对许多随机采样掩码的重建信号求平均来遵循自然的损失模拟策略,但它仅将任意顺序接口用作随机化,并且需要许多目标/参考查询。我们提出了JUMP(联合不确定性引导掩码探测),这是一种单通道评分攻击,利用了dLLMs的两个独特属性:任意顺序可解码性用于选择低参考置信度位置,并行可解码性用于通过对每个模型进行一次联合掩码查询来对所有选定位置进行评分。JUMP联合掩码选定位置并计算裁剪后的目标/参考重建差距统计量。在六个MIMIR域上的微调LLaDA - 8B - Base上,JUMP将平均ROC - AUC从SAMA的0.82提高到0.90,并显著提高了低FPR检测,同时每个目标和参考模型仅需一次选择器遍历和一次评分遍历。
英文摘要
Membership inference attacks (MIAs) test whether a candidate example was used to train a language model. Existing attacks on fine-tuned discrete diffusion language models (dLLMs) often aggregate reconstruction signals across many mask configurations, requiring repeated model evaluations. We propose Joint Uncertainty Guided Mask Probing (JUMP), an efficient MIA that exploits the ability of dLLMs to predict masked tokens in parallel. Using the pre-fine-tuning checkpoint as a reference, JUMP selects low-confidence positions, masks them jointly, and aggregates clipped target-reference reconstruction gaps. This focuses the attack on positions that reveal stronger membership signals from fine-tuning. After mask selection, all selected tokens are evaluated with one scoring query per model. Across six MIMIR domains, JUMP improves mean ROC-AUC over a prior multi-mask attack from 0.819 to 0.902 on LLaDA and from 0.851 to 0.942 on Dream. Including mask selection, it requires only three forward passes per example, compared with 32 for the baseline. We further extend JUMP to the target-only setting by replacing target-reference scoring with relative token preference, which compares the observed token with alternative predictions at the same masked position. Target-Only JUMP achieves mean ROC-AUCs of 0.609 and 0.638 on LLaDA and Dream.
发表机构
- Yonsei University(延世大学)
机构由 AI 辅助整理,请以论文原文为准。