arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合成数据中的隐藏威胁:通过良性文本注入隐蔽的针对性偏见

Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text

Minkyung Cho, Jihyo Kim, SeungWoo Song, Junghun Yuk, Minjoon Kee, Hoyun Song, KyungTae Lim

arXiv 2608.30619首次发表:更新:

发表机构

KAIST; Dankook University(韩国科学技术院; 檀国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现可通过语义良性的合成数据,利用潜意识学习机制向对齐LLM注入针对性社会偏见,同时保留模型通用任务能力,揭示了合成数据训练LLM的新安全风险,提出基于对数线性的评分或可用于筛查此类数据。

AI 中文摘要

合成数据正越来越多地用于训练大型语言模型(LLMs),但其安全影响仍鲜为人知。先前关于潜意识学习的研究表明,模型可从看似不相关的训练数据中继承行为特征。本研究探究是否可利用此类机制,通过语义上良性的合成数据向对齐模型注入针对性社会偏见。我们构建了一个流水线,其中未对齐的教师模型生成跨领域的过滤后合成数据集,涵盖创意写作、代码生成等领域,这些数据集随后用于微调对齐的学生模型。实验显示,看似良性的合成数据可作为隐蔽信道,在很大程度上保留学生模型通用任务能力的同时传递针对性偏见。这些结果揭示了合成数据驱动的LLM训练流水线中此前未被充分探索的安全风险,并强调需要改进防护措施。作为实现该目标的一个可能步骤,我们建议基于对数线性的评分或许可为筛查看似良性的合成数据提供有用信号。

英文摘要

Synthetic data is increasingly used to train large language models (LLMs), yet its security implications remain poorly understood. Prior work on subliminal learning suggests that models can inherit behavioral traits from seemingly unrelated training data. In this work, we investigate whether such mechanisms can be exploited to inject targeted social biases into aligned models through semantically benign synthetic data. We construct a pipeline in which a misaligned teacher model generates filtered synthetic datasets across domains such as creative writing and code generation, which are then used to fine-tune aligned student models. Our experiments show that benign-looking synthetic data can act as a covert channel for transmitting targeted biases while largely preserving the student model's general task capabilities. These results reveal a previously underexplored security risk in synthetic data-driven LLM training pipelines and highlight the need for improved safeguards. As one possible step toward this goal, we suggest that log-linearity-based scoring may provide a useful signal for screening seemingly benign synthetic data.

CommentsTo be published in EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑