发表机构
University of Technology Sydney; Xi’an Jiaotong University; Southeast University(悉尼科技大学; 西安交通大学; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过转向向量恢复,将阈下学习建模为噪声线性逆问题,发现软监督可少量数据精确恢复,硬监督需大量数据,因Fisher逆放大标签噪声。
AI 中文摘要
阈下学习使学生能够从语义无关的数据中继承教师的行为特征,然而已发表的演示通常需要数万个载体示例。我们探究这一数据需求的来源。我们的测试平台是阈下转向:教师特征为已知的残差流向量$\Delta_T$,因此迁移可直接作为参数恢复来衡量。在相同的载体前缀上,我们比较了词级(硬)NLL监督与全分布(软)KL监督。在初始化时,两个目标给出几乎共线的梯度,且都与$\Delta_T$对齐不佳。然而,在迭代优化下,它们出现分歧:软监督从几百个载体中几乎精确恢复$\Delta_T$,而硬监督即使有数万个载体也远低于此。我们通过将转向向量蒸馏视为噪声线性逆问题来解释这一差距。局部上,载体任务通过其Fisher矩阵$F$映射特征,因此梯度指向$F\Delta_T$而非$\Delta_T$。梯度下降随后充当$F$的逐渐减少阻尼的逆。使用软目标时,此逆恢复低曲率方向。使用硬标签时,它还会放大这些方向上的采样噪声。结果是,最优反演深度随独立载体数量增长。在Qwen2.5-7B和Gemma-2-9B上的实验证实了四个预测:初始梯度的Fisher失真、从陡峭到平坦方向的恢复顺序、随数据规模变化的最优深度,以及从固定提示池中重新采样完成与添加新提示效果相同。在此设置中,大型载体数据集的需求主要不是为了揭示特征,而是为了抑制Fisher反演放大的标签噪声。代码可在\url{此https URL}获取。
英文摘要
Subliminal learning lets a student inherit a teacher's behavioral trait from semantically unrelated data, yet published demonstrations typically require tens of thousands of carrier examples. We ask where this data requirement comes from. Our testbed is subliminal steering: the teacher trait is a known residual-stream vector $Δ_T$, so transfer can be measured directly as parameter recovery. On identical carrier prefixes, we compare token-level (hard) NLL supervision with full-distribution (soft) KL supervision. At initialization the two objectives give nearly collinear gradients, and both align poorly with $Δ_T$. Under iterative optimization, however, they diverge: soft supervision recovers $Δ_T$ almost exactly from a few hundred carriers, while hard supervision stays well below it even with tens of thousands. We explain this gap by casting steering-vector distillation as a noisy linear inverse problem. Locally, the carrier task maps the trait through its Fisher matrix $F$, so gradients point toward $FΔ_T$ rather than $Δ_T$. Gradient descent then acts as a progressively less-damped inverse of $F$. With soft targets, this inverse restores low-curvature directions. With hard labels, it also amplifies the sampling noise in those same directions. The result is an optimal inversion depth that grows with the number of independent carriers. Experiments on Qwen2.5-7B and Gemma-2-9B confirm four predictions: the Fisher distortion of the initial gradient, recovery ordered from steep to flat directions, an optimal depth that shifts with data scale, and the finding that resampling completions from a fixed prompt pool works as well as adding new prompts. In this setting, large carrier datasets are needed less to reveal the trait than to suppress label noise amplified by Fisher inversion. Code is available at \url{https://github.com/luoyuchenmlcv/subliminal-data}.