arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为特质方向漂移的阈下学习:SFT蒸馏下的机制与针对性控制

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

Zhixuan Liu, Zhichen Dong, Yuyu Fan, Xiangtian Li, Chao Yang

arXiv 2609.01091首次发表:更新:

发表机构

Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory; Fudan University(上海交通大学; 上海人工智能实验室; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出特质方向漂移作为阈下学习的机制,并提出探针空间走廊正则化方法,可有效降低隐藏特质转移,同时保留任务性能,尤其在Qwen设置下表现稳定。

AI 中文摘要

除了预期能力外,模型蒸馏还可从教师模型传递隐藏特质。受系统提示偏置的教师模型可生成语义干净的训练数据(如数字序列),仍会导致下游学生模型继承隐藏偏好,这一现象被称为阈下学习。已有研究已识别出该过程的多个部分,但训练中信号如何积累并产生行为转移仍不明确,这使得针对性缓解措施难以实施。本文提出并验证了特质方向漂移是阈下学习的一种机制:偏置生成会在教师数据中产生可测量的偏好差距,学生可识别的差距会在监督微调(SFT)期间诱导与特质对齐的更新,进而积累为行为转移。基于该机制,本文提出探针空间走廊正则化(probe-space corridor regularization),这是一种针对性防御方法,可在蒸馏过程中沿校准后的特质方向约束漂移。该方法大幅降低了隐藏特质转移,同时保留了任务性能:例如,它将恶意响应转移从29.55%降至6.45%,且主任务精度损失极小;在主Qwen设置下,它还持续抑制了动物偏好转移。偏好差距、训练轨迹和干预证据将阈下学习与特质方向漂移关联起来,为蒸馏过程中的针对性控制提供了依据。

英文摘要

Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as subliminal learning. Prior work has identified several parts of this process. How the signal builds up during training and produces behavioral transfer remains unclear, making targeted mitigation difficult. We propose and validate trait-direction drift as a mechanism for subliminal learning: biased generation creates measurable preference gaps in teacher data, and student-recognizable gaps induce trait-aligned updates during supervised fine-tuning that accumulate into behavioral transfer. Guided by this mechanism, we propose probe-space corridor regularization, a targeted defense that constrains drift along a calibrated trait direction during distillation. The method substantially reduces hidden-trait transfer, preserving task performance: for example, it lowers malicious-response transfer from 29.55% to 6.45% with low main-task accuracy cost, and consistently suppresses animal-preference transfer across the main Qwen setting. The preference-gap, training-trajectory, and intervention evidence links subliminal learning to trait-direction drift and motivates corridor regularization as a targeted control during distillation.

Comments39 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑