发表机构
Northeastern University; Eleos AI Research(东北大学; Eleos AI 研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过低秩适配器训练模型按潜在偏好决策,发现持续微调可涌现准确自我报告,并利用归因修补识别出忠实与不忠实模型的结构差异,为从机制上验证内省提供了新方法。
AI 中文摘要
大型语言模型对自身的断言既具有重大影响,又日益难以仅从行为上进行验证。我们如何区分看似合理的虚构与真正的内省?在本文中,我们在受控环境中识别了忠实自我报告(faithful self-report)的机制性特征。通过使用低秩适配器(low-rank adapters),我们训练模型根据潜在的线性偏好函数(latent linear preference functions)为虚构角色做出决策。我们发现,在隐式决策任务上进行持续微调,即使没有显式的自我报告监督,也能导致模型对其所学偏好的准确自我报告(accurate self-reporting)的出现。我们围绕这一涌现现象提出两个研究问题。第一:准确自我报告的出现是否伴随着模型中可测量的结构变化?权重消融(weight ablations)和冻结层实验(frozen-layer experiments)共同表明,在训练过程中,偏好表征(preference representations)会转移到更早的层,这与以下假设一致:忠实的自我报告要求偏好位于预先存在的语言化机制(verbalization mechanisms)可以访问的位置。第二:这些结构差异能否区分忠实模型与不忠实模型?通过使用归因修补(attribution patching),我们发现忠实模型在决策任务和自我报告任务之间表现出显著更高的归因相似性(attribution similarity)——这是忠实自我报告的一个机制性特征,无需我们理解报告本身的内容。以往关于自我报告的研究在行为层面观察到模型可能忠实或不忠实;我们的工作提出,至少在我们的受限设置中,可以通过检查网络本身的结构来区分这两种计算模式。
英文摘要
Large language models make claims about themselves that are both consequential and increasingly difficult to verify from behavior alone. How can we distinguish plausible confabulations from genuine introspection? In this paper, we identify mechanistic signatures of faithful self-report in a controlled setting. Using low-rank adapters, we train models to make decisions on behalf of fictitious characters, according to latent linear preference functions. We find sustained fine-tuning on an implicit decision task can lead to the emergence of accurate self-reporting of models' learned preferences, even without explicit self-report supervision. We ask two research questions about this emergent phenomenon. First: is the emergence of accurate self-reporting accompanied by a measurable structural change in the model? Weight ablations and frozen-layer experiments together indicate that preference representations shift to earlier layers over training, consistent with the hypothesis that faithful self-report requires preferences to be located where pre-existing verbalization mechanisms can access them. Second: can these structural differences distinguish faithful models from unfaithful ones? Using attribution patching, we find that faithful models exhibit significantly higher attribution similarity between the decision-making and self-report tasks -- a mechanistic signature of faithful self-report that does not require us to understand the content of the report itself. Previous work on self-report has observed behaviorally that models can be faithful or unfaithful; our work proposes that, at least in our restricted setting, it is possible to distinguish between the two patterns of computation by examining the structure of the networks themselves.
CommentsPublished at COLM 2026. Project page: https://iii.baulab.info