发表机构
Khoury College of Computer Sciences, Northeastern University(东北大学胡里计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示众包微调中恶意用户可通过少量投毒样本大幅放大对其他用户专有指令的提取,且现有数据过滤难以有效防御。
AI 中文摘要
监督微调(SFT)被广泛用于将大型语言模型适配到下游任务。众包用户对话是一种成熟的规模化收集SFT数据的方法,同时减少了对昂贵的人工标注的需求。然而,这也允许不可信用户向微调流程贡献数据。我们研究了由此设置产生的一个未被充分探索的隐私风险:恶意用户能否投毒一小部分众包数据,以放大对其他用户贡献的、此前未见过的指令的提取?我们证明,仅使用对部署模型的黑盒、仅输出访问即可实现这一点。在四个模型和两个数据集上的实验表明,训练数据提取显著增加:仅使用50个投毒样本,在OpenMathInstruct上对Qwen2.5-14B的近逐字提取率达到无投毒时的3.71倍,在AceReason上对Llama-3.1-8B达到3.08倍。数据过滤在检测投毒样本方面也基本无效:即使最佳方法也仅达到0.378的F-1分数,使得大多数投毒样本未被检测到。这些发现表明,看似良性的众包贡献可以放大其他记录的泄漏,同时仍然难以通过数据过滤识别。
英文摘要
Supervised fine-tuning (SFT) is widely used to adapt large language models to downstream tasks. Crowdsourcing user conversations is an established approach to collecting SFT data at scale while reducing the need for costly manual annotation. However, it also allows untrusted users to contribute data to the fine-tuning pipeline. We investigate an underexplored privacy risk arising from this setting: can a malicious user poison a small fraction of the crowdsourced data to amplify extraction of previously unseen instructions contributed by other users? We show that this is possible using only black-box, output-only access to the deployed model. Experiments across four models and two datasets demonstrate substantial increases in training-data extraction: with only 50 poisoned examples, near-verbatim extraction reaches $3.71\times$ the rate without poisoning for Qwen2.5-14B on OpenMathInstruct and $3.08\times$ for Llama-3.1-8B on AceReason. Data filtering also proves largely ineffective in detecting poisoned samples: even the best-performing method achieves only 0.378 in F-1 score, leaving the majority of poisoned samples undetected. These findings demonstrate that seemingly benign crowdsourced contributions can amplify leakage of other records while remaining difficult to identify through data filtering.