发表机构
Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory; Fudan University(上海交通大学; 上海人工智能实验室; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现扩展模型生成的蒸馏数据可使学生模型中教师的潜在特征更易被检测,且该效应跨多种场景存在,建议扩展数据时结合感知特征的整理与评估。
AI 中文摘要
扩展模型生成的数据通常被视为可提升蒸馏效果:更多示例应能增加覆盖范围、减少噪声并培养出更强的学生模型。我们揭示了第二种效应:更大的数据集能让训练后的学生模型中更易检测到细微的教师特定信号,即便示例偏离任务且从未提及该特征。在受阈下学习启发的受控设置中,被诱导表达目标特征的教师生成受限的偏离任务数据,例如仅含数字的补全内容。在不同量的独立偏离任务数据上训练的学生模型,会在单独的领域中接受评估,通过匹配无特征对照组来分离目标特定的迁移效果。我们的主要发现是,更大的独立数据集能使教师诱导的特征在学生后续行为中更清晰地凸显。其他看似合理的特征也可能随规模增强,但目标特征通常增长更多。当小规模学生模型已倾向于目标特征时,扩展数据主要放大该行为;当它倾向于相关或显著的替代特征时,更多数据可将行为转向预期的目标特征。对学习到的LoRA更新的分析显示出平行趋势。这些效应跨模型家族、特征类型、多特征设置及跨模型迁移均存在。我们的结果表明,扩展生成的蒸馏数据应与感知特征的整理和评估相结合,即便数据看似偏离任务或无害。
英文摘要
Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data, such as number-only completions. Students trained on different amounts of independent off-task data are evaluated in a separate domain, with matched no-trait controls isolating target-specific transfer. Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior. Other plausible traits may also strengthen with scale, but the target usually grows more. When the small-scale student already favors the target, scaling mainly amplifies that behavior; when it favors a related or salient alternative, more data can shift behavior toward the intended trait. Analyses of learned LoRA updates show a parallel trend. These effects appear across model families, trait types, multi-trait settings, and cross-model transfer. Our results suggest that scaling generated distillation data should be paired with trait-aware curation and evaluation, even when the data appears off-task or benign.