发表机构
CentraleSupelec; Université Paris-Saclay; IRT Saint Exupéry; Mila - Quebec AI Institute(中央高等电力学院; 巴黎萨克雷大学; 圣埃克苏佩里信息技术研究院; 米拉-魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探究语言模型训练中概念功能充分性的时间点,通过分解激活、选择软掩码等方法,比较多任务充分性并在七个模型上验证,发现下游掩码软质量少于重建掩码且预测分布偏移小。
AI 中文摘要
深入理解模型及其学习机制,需识别其内部结构何时变得有用,而非仅关注最终状态。本研究通过概念动态性展开:在每一层和每个检查点,分解激活、选择稀疏软掩码,并将掩码重建注入模型。因此,概念分析从功能层面进行测试:掩码仅在干预下保留目标时才有用。我们比较激活重建、线性可解码性、真实下游保留及学习对齐下的检查点转移的充分性。该框架将分解假设视为假设而非可解释性保证,监测跨检查点的功能充分性及学习对齐下从源到最终的可重建性。在七个模型共有的固定惩罚操作点,下游掩码保留的软质量远少于重建掩码;预测分布偏移仍较小。
英文摘要
Understanding a model and its learning mechanisms in depth requires identifying when its internal structures become useful, rather than simply looking at the final state. We study this through concept dynamics: at each layer and checkpoint, we decompose activations, select sparse soft masks, and inject masked reconstructions into the model. Concept analysis is therefore tested functionally: a mask is useful only insofar as it preserves a target under intervention. We compare sufficiency for activation reconstruction, linear decodability, true downstream preservation, and checkpoint transfer under learned alignment. The framework treats decomposition assumptions as hypotheses rather than interpretability guarantees, monitoring functional sufficiency across checkpoints and source-to-final reconstructability under learned alignment. At the shared fixed-penalty operating point across seven models, downstream masks retain substantially less soft mass than reconstruction masks; predictive-distribution shifts remain small.