arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23260cs.LGcs.AI

为何幽灵输出能教学:基于核方法的潜意识学习理解

Why Ghost Outputs Teach: A Kernel-Based Understanding of Subliminal Learning

Zhe Li, Bicheng Ying, Chaosheng Dong, Haibo Yang

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过推导链式跨任务核,从学习动力学角度解释了潜意识学习机制,证明幽灵输出监督在共享初始化下形成半正定结构,并验证了维度秩瓶颈与高熵输入的有效性。

中文摘要 AI 辅助

潜意识学习(Subliminal Learning, SL)是一种近期发现的现象,其中学生模型通过匹配教师模型中看似无关的辅助输出来获得下游任务能力,尽管从未观察过任务标签、任务特定输出或原始训练数据。虽然近期研究已识别出潜意识信号可能存在的区域,但支撑该现象的优化机制仍鲜为人知。在本工作中,我们通过学习动力学的视角提供了对SL的机制性理解。具体而言,我们推导出一个链式跨任务核,该核通过共享骨干表示将幽灵输出监督与任务预测的变化显式地联系起来。我们的统一分析框架为SL中的三个核心经验谜题提供了严格的数学解释:(i)在共享初始化下,转移算子形成严格半正定(Positive Semi-Definite, PSD)结构,保证幽灵输出优化在无显式标签暴露的情况下使学生模型与教师的真实任务目标对齐;(ii)幽灵输出的维度充当显式的秩瓶颈,控制任务相关特征的转移;(iii)合成的高熵输入充当宽带探针,最大化跨任务核重叠,解释了为何随机噪声在潜意识转移中始终优于结构化数据。在典型幽灵输出设置上的实验验证了所有三个理论预测,提供了首个基于学习动力学的理论解释,说明幽灵输出监督如何引发潜意识学习。

英文摘要

Subliminal Learning (SL) is a recently identified phenomenon in which a student model acquires downstream task capabilities by matching seemingly unrelated auxiliary outputs from a teacher, despite never observing task labels, task-specific outputs, or the original training data. While recent studies have identified where subliminal signals may reside, the optimization mechanism underlying this phenomenon remains poorly understood. In this work, we provide a mechanistic understanding of SL through the lens of learning dynamics. Specifically, we derive a chained cross-task kernel that explicitly links ghost-output supervision to changes in task predictions through shared backbone representations. Our unified analytical framework provides a rigorous mathematical explanation for three central empirical puzzles in SL: (i) under shared initialization, the transfer operator forms a strictly Positive Semi-Definite (PSD) structure, guaranteeing that ghost-output optimization aligns the student with the teacher's true task objective without explicit label exposure; (ii) the ghost-output dimensionality acts as an explicit rank bottleneck governing the transfer of task-relevant features; and (iii) synthetic, high-entropy inputs function as broadband probes that maximize cross-task kernel overlap, explaining why random noise consistently outperforms structured data for subliminal transfer. Experiments on the canonical ghost-output setting validate all three theoretical predictions, providing the first learning-dynamics-based theoretical explanation of how ghost-output supervision gives rise to subliminal learning.

发表机构

  • Rochester Institute of Technology(罗切斯特理工学院)
  • Google Inc(谷歌公司)

机构由 AI 辅助整理,请以论文原文为准。

↑