使用文本优化来言语化阈下学习效应
Verbalizing Subliminal Learning Effects Using Text Optimization
浏览论文内容
中文总结 AI 辅助
本文提出SALVE方法,利用文本优化将阈下学习效应言语化为可读提示,在多种设置下有效检测教师特征,深化对阈下学习的理解。
中文摘要 AI 辅助
阈下学习是一种现象,其中蒸馏数据集从教师模型中传递了在数据集本身中未以可读方式编码的特征。这为模型开发带来了新的挑战,并因数据投毒而产生了新的风险。在这项工作中,我们使用文本优化来检测阈下学习效应,并将其描述为可读的提示。来自带提示的教师的阈下学习激发了我们的方法。我们观察到这是上下文蒸馏的一个特例,并利用这一观察在理论上表明,带提示的阈下学习数据集能够识别教师的提示。我们将恢复该提示的问题简化为一个文本优化问题,并提出了一种近似求解的方法。我们的方法SALVE(搜索辅助的潜在言语化)优化一个软提示,查询同一模型将其言语化为文本,并使用束搜索使言语化可靠。在标准的阈下学习设置中,SALVE可靠地恢复了命名教师特征的清晰提示,而常见的文本优化方法无法做到这一点。此外,我们发现存在一些设置,即使阈下学习失败,SALVE也能从数据集中恢复教师的特征,但修改学生训练以改进上下文蒸馏可能会产生阈下学习效应。最后,我们展示了SALVE在三种额外设置中检测阈下学习效应:(1)阈下学习数据与无关数据的混合,(2)通过激活引导使教师产生偏差时生成的数据,以及(3)通过Logit-Linear选择选出的真实偏好数据子集。总的来说,我们的结果加深了对阈下学习的理解,并提出了SALVE作为一种主动检测阈下学习效应的方法。
英文摘要
Subliminal learning is a phenomenon in which a distillation dataset transmits traits from the teacher model that are not legibly encoded in the dataset itself. This introduces a new challenge for model development and creates new risks from data poisoning. In this work, we use text optimization to detect subliminal learning effects and describe them as legible prompts. Subliminal learning from a prompted teacher motivates our approach. We observe that this is a special case of context distillation and leverage this observation to show that, in theory, the prompted subliminal learning dataset identifies the teacher's prompt. We reduce recovering this prompt to a text optimization problem and present a method to approximately solve it. Our method, SALVE (Search-Aided Latent Verbalization), optimizes a soft prompt, queries the same model to verbalize it as text, and uses beam search to make the verbalization reliable. In the standard subliminal learning setting, SALVE reliably recovers legible prompts that name the teacher's trait, while common text optimization methods fail to do so. In addition, we find that there are settings in which SALVE recovers the teacher's trait from a dataset even when subliminal learning fails, but that modifying student training to improve context distillation can create subliminal learning effects. We lastly show that SALVE detects subliminal learning effects in three additional settings: (1) mixtures of subliminal learning data and unrelated data, (2) data generated when the teacher is biased via activation steering, and (3) subsets of real preference data selected via Logit-Linear Selection. Overall, our results deepen our understanding of subliminal learning and present SALVE as a method to proactively detect subliminal learning effects.
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。