arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当激活预言机学会不读取时:微调预言机中的特定概念盲点

When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles

Tobias Bersia, Tatiana Gaintseva

arXiv 2607.23379首次发表:更新:

发表机构

Queen Mary University of London(伦敦玛丽女王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在禁忌词猜测设置下,激活预言机经微调后成为特定概念反读取器的现象,发现其无法恢复自身训练时存在的概念,失败源于读取路径,揭示了行为泄漏等因素可分离,引发对可解释性接口可靠性的担忧。

AI 中文摘要

激活预言机(AO)是经过训练以回答有关另一个模型内部激活的自然语言问题的语言模型。它为从模型状态中读取隐藏信息提供了灵活接口。然而,AO本身是学习系统,其答案受训练数据、目标和学习报告行为影响。我们在禁忌词猜测设置中研究此问题,在该设置中主题模型经微调在内部使用隐藏概念并避免直接披露。与预期相反,经微调的AO会变成特定概念的反读取器,它们选择性地无法恢复在自身训练期间持续存在的概念。这种失败不能简单地用主题或预言机表示中不存在该概念来解释,目标在预言机内部仍可解码,而LogitLens和层消融分析表明失败出现在AO读取路径中。我们的结果表明行为泄漏、表示级可解码性和AO可表达性可能会分离,这引发了对学习到的可解释性接口可靠性的担忧。

英文摘要

Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hidden information from model states, especially when relevant information is internally represented but absent or incomplete in visible behavior. However, AOs are themselves learned systems: their answers are shaped by training data, objectives, and learned reporting behavior, rather than being neutral readouts of represented information. We study this in a controlled Taboo Word Guessing setting, where subject models are fine-tuned to internally use a hidden concept while avoiding direct disclosure. Contrary to the expectation that an AO trained on such a subject becomes a specialist reader, we find that fine-tuned AOs can become concept-specific anti-readers: they selectively fail to recover the concept persistently present during their own training. This failure is not simply explained by absence of the concept from the subject or oracle representations: the target remains decodable inside the oracle, while LogitLens and layer-ablation analyses indicate that the failure arises in the AO readout pathway. Our results show that behavioral leakage, representation-level decodability, and AO-verbalizability can come apart, raising a reliability concern for learned interpretability interfaces.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑