训练模型,而非读者:可验证激活解释的可解码性监督
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
浏览论文内容
中文总结 AI 辅助
研究自然语言自动编码器对隐藏激活解释评分的问题,提出RECAP等方法,能使指定内容可解码,提高解释忠实度与安全性,如在预训练模型上实现可靠探测解码,提升对真实声明评分及识别谎言能力。
中文摘要 AI 辅助
自然语言自动编码器通过重建对隐藏激活的解释进行评分:如果可以从激活中再生,则该解释被视为忠实的。该测试在结构上对个别错误声明不敏感:如果翻转一个声明不会改变重建,则该声明永远不会受到惩罚。我们表明该测试可以通过两种方式通过,但都不忠实。在发布的Qwen-2.5-7B发声器上,解释的重建远高于随机水平,而约2%的特定声明依赖于重建,因此分数跟踪要点而非特定事实。在精确的合成地面真值下,标准方法在5次运行中有5次开发了共同适应的私有代码(重建依赖的错误措辞),并且不改变目标模型的修复方法无济于事。我们贡献了两种审计协议,即基于事实与真实的交叉和评估器交换,以及RECAP(通过共同训练的辅助预测器进行可读编码):与目标模型一起训练的线性头,以保持指定内容的可解码性。在经过RECAP训练的沙盒模型上,新的发声器真实地陈述了指定内容,并且代码消失,成本为+0.001纳特。这在预训练的Pythia-160M上得到了复制:内容变得可靠地可探测解码,尽管新的发声器只能部分传达它(真实率为0.44-0.46,而对照接近零)。对于可解释性,高重建并不能证明个别声明。对于人工智能安全,RECAP使指定的内部内容可以独立于模型可能操纵的散文进行检查:独立探针对发声器的真实声明的评分高于错误声明(AUC为0.96,而没有RECAP时为0.82)。针对编辑解释以在说谎时最大化重建分数(抑制约87%的说谎惩罚)的对手,RECAP探针仍然可以标记谎言(AUC为0.95),而对照探针则降至随机水平(0.51)。
英文摘要
Natural-language autoencoders score explanations of hidden activations by reconstruction. An explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims. If flipping a claim does not change the reconstruction, the claim is never penalized. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are ones the reconstruction depends on, so the score tracks gist, not specific facts. Under exact synthetic ground truth, standard training consistently develops co-adapted private codes (false wording the reconstruction depends on), and fixes that leave the target model unchanged do not help. We contribute two audit protocols, the comparison of grounding and truth and the swap to an independent evaluator, and RECAP (Readable Encodings via Co-trained Auxiliary Predictors), linear heads trained alongside the target model to keep designated content decodable. On RECAP-trained sandbox models, fresh verbalizers state the designated content truly and the codes vanish, at a +0.001-nat cost. This replicates on a pretrained Pythia-160M. The content becomes reliably decodable by a probe, though a fresh verbalizer conveys it only in part (truth 0.44-0.46 vs a near-zero control). For interpretability, high reconstruction does not certify individual claims. For AI safety, RECAP makes designated content checkable against a probe rather than asserted by prose a model can game. An independent probe ranks the verbalizer's true claims above its false ones (AUC 0.96 vs 0.82 without RECAP). Against an adversary that edits an explanation to maximize the score while lying (suppressing ~87% of its lie penalty), the RECAP probe still flags the lies (AUC 0.95) while the control probe collapses to chance (0.51).